mirror of
https://github.com/apache/nuttx.git
synced 2026-10-10 07:40:27 +00:00
An unprivileged task that touched memory it does not own took the whole
system down. The MMU refused the access, as it should, and then
x86_64_fault_panic_isr() panicked -- so a contained application bug
became a system-wide outage. x86_64 had no user-fault recovery at all,
while arm64, RISC-V and esp32s3 each have one, and ISR13 and ISR14 here
went straight to a handler whose own comment says "Don't even brother
to recover, just dump the regs and PANIC."
The decision has to be made in the handler. The user-task check in
_assert() looks like it already covers this, but a fault arrives as an
exception, so up_interrupt_context() is already true by the time it is
reached and the panic branch is taken no matter who faulted.
The CPL in the saved CS is the whole test, and on this architecture it
is exact: everything that runs on behalf of a user task inside the
kernel -- a system call body, an interrupt handler, a kernel thread --
runs in ring 0, so a fault there is correctly refused recovery. That is
what RISC-V reads out of STATUS_PPP and arm64 out of SPSR_MODE_EL0T.
x86_64 cannot use TCB_FLAG_SYSCALL the way those two do: it has never
set it, and making it do so is a separate change with its own hazards
(see the note in x86_64_syscall()). The task type is checked as well,
as RISC-V does, so that a frame that cannot be trusted -- an early-boot
fault, before any user task exists -- cannot talk its way in with a
stale selector.
Recovery is the same shape as the other architectures -- RIP to _exit,
first argument SIGSEGV, TCB_FLAG_FORCED_CANCEL raised -- plus the three
x86_64 specifics:
* CS and SS move to the kernel selectors together with RIP. The frame
being rewritten is the one x86_64_fullcontextrestore() is about to
iretq from, and iretq takes the target privilege level from the CS
it pops; in long mode it pops SS with it even when the level does
not change.
* RSP moves to the top of the task's own kernel stack. _exit() must
not run on the user stack the fault came from, because the address
environment that stack belongs to is torn down while _exit() is
still running. The kernel stack is free -- the fault was taken in
user mode, so no system call of this task is in flight. The -8
reproduces the offset a call would have left, which is what the SysV
ABI states its 16-byte rule against and what up_initial_state() sets
up for the same reason.
* RFLAGS is reset to the value up_initial_state() gives a new thread
rather than carried over. The faulting task's flags are its own to
set, and DF in particular must be clear on entry to any C function.
ISR6 is routed through the same handler. An invalid opcode is
attributable to the instruction that raised it, so a user task running
garbage should die on its own rather than take the system with it --
the same conclusion esp32s3 reached for EXCCAUSE_ILLEGAL. ISR8 keeps
panicking unconditionally: a double fault says an exception could not be
delivered at all, and there is nothing left to trust.
The message gives the fault address (CR2) only for a page fault, because
the other exceptions do not set CR2.
X86_GDT_PL_MASK and X86_GDT_RPL_USER now live in intel64/arch.h beside
the selectors they mask; x86_64_fork.c had a private copy of the latter.
Verified on qemu-intel64:knsh_romfs under QEMU TCG with
apps/examples/sandbox. The probe targets kernel .text at _stext
(0x100909000). This config sets CONFIG_RAM_START to 0x0, a Kconfig
default, so the address is given on the command line. The "x" probes
call the address; that mode was added to the sandbox locally for the
test.
sandbox r 0x100909000 -> Exception 14, error code 5, 2816
sandbox w 0x100909000 -> Exception 14, error code 7, 2816
sandbox x 0x100909000 -> Exception 14 at RIP=100909000, 2816
sandbox r 0x8000000000000000 -> Exception 13 (non-canonical), 2816
sandbox x 0x80000000d -> Exception 6 at RIP=80000000d, 2816
All five probes run in one boot and report CONTAINED. The offender
dies with SIGSEGV, the parent survives, a canary thread keeps running,
and the memory and descriptors of the offender come back. The shell
answers after the last probe. Without this change each probe panics
the system.
ostest exits with status 0 on knsh_romfs, fork and vfork included. It
also exits with status 0 on the flat qemu-intel64:nsh with
CONFIG_SCHED_THREAD_LOCAL disabled. With it enabled, ostest faults in
sched_thread_local_test() with and without this change.
Each vector is reached deliberately. A non-canonical address is a #GP
rather than a #PF, which the read and call probes never produce on their
own. The #UD needs no corrupt binary either: the crt0 stub this
architecture links into every user ELF already ends in a ud2 at
_stext+0xd, and with CONFIG_ARCH_TEXT_VBASE at 0x800000000 the call
probe can simply be aimed at it, executing an invalid opcode in ring 3
out of the process's own text.
The negative direction was checked too, on the flat build, where the
same non-canonical read is taken at CPL 0: it reaches
x86_64_fault_panic_isr() exactly as before and panics with a full
register dump reporting CPL 0.
Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
|
||
|---|---|---|
| .. | ||
| arm | ||
| arm64 | ||
| avr | ||
| ceva | ||
| dummy | ||
| hc | ||
| mips | ||
| misoc | ||
| or1k | ||
| renesas | ||
| risc-v | ||
| sim | ||
| sparc | ||
| tricore | ||
| x86 | ||
| x86_64 | ||
| xtensa | ||
| z16 | ||
| z80 | ||
| CMakeLists.txt | ||
| Kconfig | ||