ARCH_INTEL64_HPET_ALARM is one of the three members of the "System Timer
Source" choice, but selecting it does not build. Taking qemu-intel64:nsh on
master and moving the choice off ARCH_INTEL64_TSC_DEADLINE onto it:
intel64/intel64_hpet_alarm.c:41:24: error:
'CONFIG_ARCH_INTEL64_HPET_ALARM_CHAN' undeclared
ARCH_INTEL64_HPET_ALARM_CHAN lives inside "if INTEL64_HPET" and nothing
selects INTEL64_HPET. Enabling that by hand moves the failure to link time,
because intel64_oneshot_lower.c is built only when INTEL64_ONESHOT is set:
undefined reference to `oneshot_initialize'
Enabling INTEL64_ONESHOT as well finally reaches the real problem.
intel64_oneshot_lower.c implements the counter flavour of struct
oneshot_operations_s, which exists only with ONESHOT_COUNT:
intel64_oneshot_lower.c: error: 'const struct oneshot_operations_s' has no
member named 'start_absolute'
intel64_oneshot_lower.c: error: implicit declaration of function
'oneshot_count_init'
intel64_oneshot_lower.c: error: initialization of
'int (*)(struct oneshot_lowerhalf_s *, const struct timespec *)' from
incompatible pointer type ... (four more of these)
ONESHOT, ONESHOT_COUNT and ONESHOT_FAST_DIVISION were selected by
ARCH_INTEL64_TSC_DEADLINE and by nothing else, so the other members of the
same choice could never be built.
Select the four from ARCH_INTEL64_HPET_ALARM as well. INTEL64_ONESHOT
selects INTEL64_HPET in turn, which is what brings
ARCH_INTEL64_HPET_ALARM_CHAN into existence, so one added select closes all
three stages.
Impact: build only, and only for a configuration that could not be built
before. No existing defconfig selects ARCH_INTEL64_HPET_ALARM.
Assisted-by: Claude:claude-opus-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
The reference manual (section 3.2.1) states the user has to
read GLERR and INTERR registers to clear their bits and release
ERR pin after the startup sequence. Error bits are set to 1 after
the startup if external power supply is used.
Not clearing the bits leads to subsequent read call errors if VBB
errors are checked.
Signed-off-by: Michal Lenc <michallenc@seznam.cz>
During the Toybox port to NuttX, Claude noticed that changes in the
menuconfig weren't taking affect. This issue exists for a long time on
NuttX, in fact BayLibre's presentation from 2017 make jokes about our
building system not been reliable:
https://www.youtube.com/watch?v=XUJK2htXxKw&t=320s
Stale archive members from $(AR)'s additive-only behavior can linger
after Kconfig toggles change which files provide a symbol, causing dead
weight or "multiple definition" link errors on incremental builds.
Fixed by splitting ARCHIVE into two macros: ARCHIVE keeps the original
additive behavior for apps/libapps.a, which many independent
subdirectories contribute to across a build, while the new
ARCHIVE_REBUILD deletes then archives for the far more common case
of a single Makefile building its own self-contained $(OBJS)
- all 39 such call sites now use it.
Assisted-By: Claude Sonnet 5
Signed-off-by: Alan C. Assis <acassis@gmail.com>
syslog_write_foreach() compares an unsigned count against a signed
accumulator:
size_t nwritten = 0;
ssize_t nwritten_max = -EIO;
...
if (nwritten > nwritten_max)
{
nwritten_max = nwritten;
}
return nwritten_max;
The usual arithmetic conversions promote nwritten_max to size_t, so -EIO
becomes 4294967291 on a 32-bit target, and the comparison is never true.
nwritten_max keeps its initial value and the function returns -EIO no
matter how many bytes actually went out. Observed under gdb on a running
target: nwritten == 64, nwritten_max == -5, (nwritten > nwritten_max) == 0.
Most callers discard the result -- syslog() itself returns void -- so this
is normally invisible. It becomes fatal when /dev/console is backed by
syslog_console_write(), because then stdio acts on it.
lib_fflush_unlocked() sees a negative return, sets __FS_FLAG_ERROR and
returns early, before resetting fs_bufpos. The bytes have already been
emitted, but the buffer is never cleared, so every subsequent stdio call
re-flushes the same CONFIG_STDIO_BUFFER_SIZE bytes. The console fills
with one repeated fragment and the system makes no further progress.
Reaching that state needs CONFIG_CONSOLE_SYSLOG=y together with no driver
claiming /dev/console ahead of syslog_console_init(). Three in-tree
defconfigs qualify: x86/qemu-i486:ostest, renesas/skp16c26:ostest and
x86_64/qemu-intel64:earlyfb. The other 56 CONSOLE_SYSLOG configurations
have a serial console that registers /dev/console first, which is why this
has gone unnoticed.
Introduced by 1685e8ff7b ("syslog: avoid an infinite loop if one channel
fails"), which changed nwritten_max from size_t to ssize_t = -EIO so that
an all-channels-failed case could be reported. Give nwritten the same type
so the comparison is signed, which preserves that intent: nwritten_max
stays -EIO only when no channel wrote anything. nwritten is never negative,
so the remaining comparisons against buflen are unaffected.
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
mkallsyms.py prints "Please execute the following command to install
dependencies: pip install pyelftools cxxfilt" and then calls
os._exit(errno.EINVAL). os._exit() terminates without flushing stdio, so
when stdout is a pipe -- which it always is under make -- the message sits in
the buffer and is discarded. What the developer sees is:
make[1]: *** [Makefile:65: nuttx] Error 22
with no indication of the cause anywhere in the build output. Error 22 is
just errno.EINVAL leaking out as an exit status. The same applies to
usage(), which exits ENOENT the same way.
Use sys.exit() in both places, which raises SystemExit and lets the
interpreter flush on the way out. Nothing else changes: the exit statuses
are the same, and os is no longer referenced, so the import goes too.
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
tcp_start_monitor() is called from accept() (net/inet/inet_sockif.c) for
each newly accepted connection. When the peer had already closed the
connection before accept() ran, the monitor takes an early-return path so
that any read-ahead data buffered on the connection can still be drained;
it returns OK in that case. accept() (net/socket/accept.c) then marks the
new socket _SF_CONNECTED unconditionally.
If the peer aborts the connection with an RST immediately after the
three-way handshake completes (for example any close with SO_LINGER
{1, 0}), the connection is moved to TCP_CLOSED with no buffered data, yet
accept() still hands back a socket that reports _SS_ISCONNECTED. A
subsequent blocking send() on that socket passes the connected check,
registers a send callback and waits on its semaphore forever: the only
TCP_ABORT event was delivered before the callback existed, and no further
ACK, POLL or disconnect event is generated for a closed connection, so the
waiter is never woken.
Any server that writes before reading can hit this; the telnet daemon
(netutils/telnetd) is one example, where the accepted session task blocks
in send() and never completes.
Only return OK from the already-closed path when there is actually
read-ahead data to drain. Otherwise the connection is dead, so fall
through to the -ENOTCONN return: accept() then fails cleanly instead of
handing back a socket wedged on a connection that will never make progress.
The graceful-close-with-pending-data case (the reason the OK path exists)
is preserved by the conn->readahead check.
Signed-off-by: Ricard Rosson <ricard@groundbits.com>
Assisted-by: Claude (Anthropic Claude Code)
In a protected build arm_svcall.c treats the caller's CONTROL as part of
the saved system call state: it stores it in xcp.syscall[].ctrlreturn on
entry and restores it from there on SYS_syscall_return. All three
Cortex-M profiles do this -- armv6-m, armv7-m and armv8-m each define the
field in arch/arm/include/<arch>/irq.h and use it symmetrically.
arm_fork_direct() copied sysreturn and excreturn to the child but not
ctrlreturn. The child's TCB comes from kmm_zalloc(), so the field was
zero, and CONTROL == 0 is nPRIV clear: the child returned to user space
privileged while its parent returned unprivileged. The child ran out its
life with the MPU restrictions its parent is under silently lifted, which
is the isolation BUILD_PROTECTED exists to provide.
Nothing faults, and that is why this survived. CONTROL == 0 also selects
MSP, which sounds like it should crash immediately, but NuttX already
runs Cortex-M threads on MSP -- the parent's saved value is 0x1, nPRIV
set and SPSEL clear -- so the two differ only in the privilege bit and
there is no stack change to trip over. Privileged code then passes every
test unprivileged code passes, so ostest cannot see it either.
Measured on an RP2350 (Cortex-M33) in BUILD_PROTECTED, breaking at the
nxtask_start_fork() call in arm_fork_direct() during task_fork_test:
parent ctrlreturn 0x00000001, child ctrlreturn 0x00000000. With this
change both read 0x00000001.
BUILD_FLAT is unaffected: without CONFIG_LIB_SYSCALL, nsyscalls is 0 and
the whole block is skipped. armv7-a and armv7-r are unaffected too; they
carry cpsr instead, and that is already copied.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
The `to->preemp_start = current` assignment when `to->lockcount > 0`
was executed twice under CONFIG_SCHED_CRITMONITOR_MAXTIME_PREEMPTION,
once in the main preemption block and again after the csection block.
Remove the redundant duplicate.
Signed-off-by: yushuailong <yyyusl@qq.com>
The compiler-rt builtins build globs every arm/*.S source, but several of
those hand-written assembly files require FPU features the target may not
have. Upstream compiler-rt selects them conditionally; NuttX did not, so
BUILTIN_COMPILER_RT builds for single-precision-FPU Arm targets (e.g.
Cortex-M33, -mfpu=fpv5-sp-d16) failed to assemble with errors such as
"selected FPU does not support instruction -- vadd.f64".
Filter the source list to match the configured FPU, in both the Makefile
and CMake builds:
- chkstk.S / chkstk2.S are Windows/MinGW-only stack probes, always dropped;
- with no hardware FPU (!CONFIG_ARCH_FPU) all arm/*vfp.S are dropped;
- with a single-precision FPU (!CONFIG_ARCH_DPFPU) the double-precision
*df*vfp.S routines are dropped.
Reproduced against compiler-rt 17.0.1 with arm-none-eabi-gcc 14.2 using the
Cortex-M33 single-precision flags: 18 of 86 arm/*.S files failed to assemble
(17 double-precision *df*vfp.S plus chkstk.S); after the filter all remaining
68 files assemble cleanly.
Fixes: https://github.com/apache/nuttx/issues/17386
Generated-by: Claude (Anthropic)
Signed-off-by: Udit Jain <uditjainstjis@gmail.com>
on_hci() ran the host upcall with g_sdc_dev.lock held. On nrf53 this
deadlocks the BLE link: the upcall waits for the app core,
which cannot answer while bt_hci_send() blocks on the same lock.
nrf52 modified for consistency, deadlock is not possible there.
Signed-off-by: raiden00pl <raiden00@railab.me>
Assisted-by: Claude Code
The SSE-200 subsystem the AN521 image implements lists the UART interrupts
receive first: UART0 RX is external interrupt 32 and UART0 TX is 33, which
in NuttX numbering are 48 and 49. The configuration had those two the wrong
way round, so the TX interrupt was dispatched to uart_cmsdk_rx_interrupt.
That handler clears only UART_INTSTATUS_RX, so the asserted TX status bit
survived the acknowledgement and the interrupt re-fired immediately. The
board live-locked in the interrupt handler as soon as the console emitted its
first character: the NSH banner appeared, and nothing ran afterwards, console
input included.
Note that the overflow interrupt already sits at 63, external 47, which is
where the SSE-200 map puts it only if the pair below it is receive first --
the two corrected values are the ones consistent with it. mps2-an500 already
follows the same order with RX at 16 and TX at 17.
Assisted-by: Claude Code:claude-opus-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
The rp23xx hardware headers define a register address macro for every
register, then a block of register bit definitions. In three headers a bit
definition reuses the name of a register address macro, so the register
address is silently redefined as a bit mask:
RP23XX_POWMAN_BADPASSWD address 0x40100000 -> (1 << 0)
RP23XX_POWMAN_BOD_CTRL address 0x40100018 -> (1 << 12)
RP23XX_POWMAN_DBG_PWRCFG address 0x401000a4 -> (1 << 0)
RP23XX_BUSCTRL_BUS_PRIORITY_ACK address 0x40068004 -> (1 << 0)
RP23XX_BUSCTRL_PERFCTR_EN address 0x40068008 -> (1 << 0)
RP23XX_PADS_QSPI_VOLTAGE_SELECT address 0x40040000 -> (1 << 0)
None of these headers has an in-tree user yet, which is why this has gone
unnoticed; each clash appears as a "macro redefined" warning as soon as a
driver includes the header. Code that included one of them and used the
register by name would have dereferenced 1 or 0x1000 instead of the register.
Two of the POWMAN clashes were plain duplicates. Per the RP2350 datasheet
BOD_CTRL bit 12 is ISOLATE and DBG_PWRCFG bit 0 is IGNORE, and the correctly
named RP23XX_POWMAN_BOD_CTRL_ISOLATE and RP23XX_POWMAN_DBG_PWRCFG_IGNORE were
already defined with the same values on the following lines, so the bare names
are simply removed. The blank line separating the VREG_LP_EXIT and BOD_CTRL
groups is restored at the same time; its absence is what let the duplicate
hide inside the preceding group.
The other four are single field registers whose field carries no separate name
(the datasheet and the SDK describe each as a one bit register), so their bit
definitions are renamed to <REGISTER>_MASK, following the _MASK spelling these
headers already use for a field extent, and written in hex like their peers.
The rp23xx-rv copies of the three headers are identical to the arm ones and
carry the same clashes, so they get the same change and stay in sync.
No functional change: none of the six names has any user in the tree.
Assisted-by: Claude Code:claude-opus-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
The fix was ported from the STM32G0 to all the STM32 platforms,
as the code is mostly the same hence presents the same failure
Signed-off-by: Javier Alonso <javieralonso@geotab.com>
When the interrupts get disabled, the callback(s) are still attached.
That structure is never cleared, causing several calls to attach/detach
to eventually fail as the callback queue gets full.
When the error occurs, the registration fails with error 12 (ENOMEM).
By detaching the IRQ and clearing the callbacks, this error doesn't
happen again
Signed-off-by: Javier Alonso <javieralonso@geotab.com>
The rp23xx common board sources reference rp23xx_st7735.c from both
Make.defs and CMakeLists.txt under CONFIG_LCD_ST7735, but the file was
never added. Enabling CONFIG_LCD_ST7735 on any rp23xx board therefore
fails the build on the missing source.
Add the file, modelled on the existing RP2040 sibling
boards/arm/rp2040/common/src/rp2040_st7735.c and retargeted to
rp23xx/SPI1. It implements board_lcd_initialize(), board_lcd_getdev()
and board_lcd_uninitialize(): bring up SPI1, claim the D/C (shared with
the unused SPI1 RX pad, per the rp23xx common convention), RST and BL
pads as GPIO, pulse the panel reset and bind the ST7735 driver.
Validated on silicon on a Waveshare RP2350-LCD-0.96 (RP2350A + ST7735S
160x80 IPS): the panel powers up and displays correctly, painted at boot
with no console interaction. Build-tested raspberrypi-pico-2:nsh with
CONFIG_LCD_ST7735=y (compiles and links).
Assisted-by: Claude (Anthropic Claude Code)
Signed-off-by: Ricard Rosson <ricard@groundbits.com>
A task does not necessarily own an address environment. tcb->addrenv_own is
set only by addrenv_attach(), which is reached only from addrenv_allocate();
a kernel thread never allocates one, and in a protected build nothing does --
there is a single address space for the whole system and the architecture's
up_addrenv_*() are stubs. addrenv_own is then NULL for every task, always.
That a task may have no address environment is already an expected state.
addrenv_switch() returns OK when tcb->addrenv_curr is NULL and addrenv_drop()
returns early, and every caller of addrenv_select() checks addrenv_own != NULL
before calling in: nxsched_get_stateinfo(), nxtask_argvstr(), proc_groupenv()
and the arm, arm64, risc-v and tricore up_check_tcbstack().
addrenv_take() and addrenv_give() are the only two that dereference
unconditionally. addrenv_join() calls addrenv_take(ptcb->addrenv_own) without
a check, so pthread_create() faults on &((struct addrenv_s *)NULL)->refs
whenever the calling task has no address environment. With
CONFIG_DEBUG_ASSERTIONS off the same access silently corrupts low memory
instead.
Handle NULL in both, the way the rest of the file already does.
addrenv_give() returns a non-zero count for the NULL case so that callers
never conclude an absent address environment has become unreferenced and
should be destroyed.
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
When enabling CONFIG_READLINE_EDIT the CI fails because it reports
there is not left space in this device.
Signed-off-by: Alan C. Assis <acassis@gmail.com>
When loading a zoneinfo file fails, parse a non-colon-prefixed TZ value
as a POSIX timezone string. Treat a successful tzparse() result as
success while preserving the leading-colon file-only behavior.
Assisted-by: Codex:gpt-5.6 Sol
Signed-off-by: hanzhijian <hanzhijian@zepp.com>
CMake validates a compiler by building and linking a test program. For a
bare metal cross toolchain that link cannot succeed on its own terms, and
the flags NuttX supplies for it assume a toolchain shipped with newlib:
gcc.cmake sets
set(CMAKE_EXE_LINKER_FLAGS_INIT "--specs=nosys.specs")
A perfectly usable arm-none-eabi-gcc built without those specs therefore
fails configuration before a single NuttX source file is considered:
arm-none-eabi-gcc: fatal error: cannot read spec file 'nosys.specs':
No such file or directory
CMake Error: ... CMake will not be able to correctly generate this
project.
Setting CMAKE_TRY_COMPILE_TARGET_TYPE to STATIC_LIBRARY makes the
detection step compile without linking, which is the documented approach
for cross compiling to a bare metal target. Toolchains that do ship
nosys.specs are unaffected: the flag only applies to CMake's own
detection, not to the NuttX link.
This is not arch specific, so it is set once in the top level CMakeLists.txt
alongside the toolchain-file selection, before project() triggers detection.
sim is excluded: it is hosted and links real host executables, so it keeps
the usual executable-based detection. try_compile() reads the variable from
the calling scope, so it takes effect without living in a toolchain file;
tools/toolchain.cmake.export already sets the same for the exported build.
Assisted-by: Claude Code:claude-opus-4-8
Assisted-by: Claude Code:claude-fable-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
The aligned direct-DMA receive path only invalidated the destination
buffer before the transfer in stm32_dmarecvsetup(). On the Cortex-M7
the cache can speculatively prefetch into that cacheable buffer
between the pre-DMA invalidate and DMA completion, leaving stale
lines that shadow the data just written by the IDMA, so the CPU
reads a previously cached sector instead of the freshly received
data.
Invalidate again in stm32_recvdma() once the aligned transfer
completes, before the buffer is consumed. The buffer and length are
cache-line aligned on this path, so no adjacent memory is affected.
This matches the STM32 AN4839 guidance that a cache invalidate is
required after DMA completion and before the CPU reads the updated
region, not only before the transfer starts. A related instance of
the same "invalidate too early" defect on STM32H7 SPI DMA is tracked
in apache/nuttx#11594.
Root-caused on a PX4 FMUv6C (STM32H743) board where MAVLink ULog
downloads were intermittently corrupted: forensic diffing showed
corrupted windows were exactly 32 bytes (the D-cache line size),
cache-line aligned, and byte-for-byte equal to the previous 512-byte
SD sector cached in the FAT single-sector buffer. Disabling the
D-cache made the corruption disappear, isolating the defect to cache
coherency. After this fix, downloaded files matched the source file
byte-for-byte (sha256 identical) across a 5.8 MB log spanning
thousands of sectors.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Yang-Rui Li <yang77567789@gmail.com>
Two sequencing problems in the MMC wide bus path break eMMC 4-bit
operation on hosts that program the bus width in the SDIO widebus /
clock callbacks (e.g. STM32H7):
1. The SWITCH command (CMD6) is issued before the host has switched
to wide bus operation. When the card completes the switch while
the host is still in 1-bit mode the switch never takes effect and
every following data transfer times out. Switch the host to wide
bus operation before issuing CMD6.
2. The transfer clock is selected only at the end of mmcsd_widebus(),
so the whole switch sequence runs at ID-mode clock and, on the
affected hosts, the final clock update does not take effect either,
leaving the bus at ~400 kHz. Select the MMC transfer clock before
calling mmcsd_widebus(), and pick CLOCK_MMC_TRANSFER_4BIT when wide
bus operation is active (mirroring the SD card path) so a later
clock selection cannot revert the host to 1-bit.
No behavior change for SD cards, and no change on hosts whose widebus
callback only records the requested state.
Tested on a custom STM32H743 board with eMMC: sd_bench sequential
write ~4.1 MB/s, sequential read ~6.3 MB/s (previously all data
transfers timed out).
Signed-off-by: DuoYuWang <thirteenking.wang@gmail.com>
Building the hello_d example with LDC ImportC fails because ImportC
does not correctly handle direct __uint128_t C-style casts such as
(__uint128_t)a and (__uint128_t)1.
Replace the direct casts with equivalent typed temporaries, preserving
the existing behavior while allowing hello_d to build successfully with
LDC ImportC.
Verified on sim:nsh with CONFIG_EXAMPLES_HELLO_D=y:
nsh> hello_d
Hello World, [skylake]!
DHelloWorld.HelloWorld: Hello, World!!
Signed-off-by: Ansh Rai <anshrai331@gmail.com>
dump_stacks() only checked kernelstack_sp != 0 || force, so for tasks
with no kernel stack (e.g. idle, kernelstack_base == 0) the force path
passed base 0 to dump_stackinfo() and dumped raw memory from address 0,
triggering a secondary fault that truncated the panic log. Guard with
kernelstack_base != 0.
Signed-off-by: liang.huang <liang.huang@houmo.ai>
A crash was observed when running ps: a BusFault in nxtask_argvstr
dereferencing tl_argv, because a thread's stack overflow had silently
corrupted the TLS region where tl_argv resides.
On ARMv8-M with CONFIG_ARMV8M_STACKCHECK_HARDWARE, PSPLIM was set to
stack_alloc_ptr -- the bottom of the allocation where TLS begins. The
stack grows downward and TLS occupies [stack_alloc_ptr, stack_alloc_ptr
+ tls_info_size()), so an overflow crossed into TLS and clobbered
tl_argv before SP reached the limit, going undetected until code that
read the corrupted TLS data (such as ps) hit the bad pointer.
Set the limit to stack_alloc_ptr + tls_info_size() -- the top of the
TLS region and the usable stack base -- so an overflow faults at the
TLS boundary, before any TLS byte is touched. Include <tls/tls.h>
for the tls_info_size() macro, which is the value sched reserves for the
TLS region via up_stack_frame().
Signed-off-by: Junbo Zheng <zhengjunbo1@xiaomi.com>
The remote CPU's IPI handler wrote into the caller's buffer directly,
which may not be reachable from the target CPU's address environment
under CONFIG_ARCH_ADDRENV.
Signed-off-by: liang.huang <liang.huang@houmo.ai>
Building Zig applications on sim:x86_64 with Zig 0.13.0 fails because
the generated target "x86_64-freestanding-sysv" is not recognized by
Zig. Introduce a Zig-specific ABI mapping that translates "sysv" to
"gnu" without affecting other toolchains.
After correcting the target ABI, linking still fails due to unresolved
references to __zig_probe_stack. Add -fcompiler-rt so Zig includes the
required compiler runtime during linking.
Verified on sim:nsh with CONFIG_EXAMPLES_HELLO_ZIG=y:
nsh> hello_zig
[sim]: Hello, Zig!
Fixes#19475
Signed-off-by: Ansh Rai <anshrai331@gmail.com>
If user passes NULL as buffer, the driver may crash. This is problematic
for NuttX protected and kernel builds.
Details in the issue: https://github.com/apache/nuttx/issues/19473
While this fixes the issue with a typical NULL pointer, fundamentally this
will be addressed with the implementation of access_ok().
https://man7.org/linux/man-pages/man2/access.2.html
Signed-off-by: Catalin Visinescu <catalin_visinescu@yahoo.com>
riscv_fault_handler() only checked the task type and SYSCALL flag, so a
fault taken inside an interrupt handler (running in kernel mode) was
blamed on the user task that happened to be interrupted and killed with
SIGSEGV, hiding the real kernel bug. Use the STATUS_PPP bit of the trap
frame to tell whether the fault originated from user mode: only then is
it safe to kill the task.
Signed-off-by: liang.huang <liang.huang@houmo.ai>
riscv_fillpage() unconditionally panics the kernel on an unmappable
or invalid access, even when the fault is caused by a user task. A
user-space fault should terminate that task with SIGSEGV instead of
taking down the whole system, matching the existing behavior for
other exceptions in riscv_exception().
Signed-off-by: liang.huang <liang.huang@houmo.ai>
riscv_fillpage() is the LOADPF/STOREPF handler used under
CONFIG_PAGING. It checked whether intermediate page table levels were
already allocated, but never checked the final leaf PTE before
installing a new mapping.
RISC-V raises the same LOADPF/STOREPF cause both when a leaf PTE is
absent (a real fault) and when it is present but its permission bits
don't satisfy the access, e.g. a store to a .text page whose write
access was revoked after ELF loading. The two cases are
indistinguishable from mcause alone.
Treating both cases as "page missing" let riscv_fillpage silently
allocate a fresh, zeroed physical page over an existing mapping,
discarding the old page (a leak) and defeating whatever permission
that mapping was enforcing. Reproduced on real hardware: a user-space
store to an already-loaded .text page got a fresh writable page
instead of being rejected.
Check the leaf PTE's valid bit before allocating; if a mapping already
exists, panic instead of overwriting it.
Signed-off-by: liang.huang <liang.huang@houmo.ai>
Two related defects corrupt CDC-NCM transmit once TCP write buffers make TX
bursty (a single txavail poll drains many queued segments back-to-back through
cdcncm_send):
1. Buffer-reuse race. cdcncm coalesces datagrams into the single pre-allocated
wrreq->buf that the USB controller transmits directly from, but cdcncm_send
formatted a new NTB batch into it (cdcncm_transmit_format) without first
waiting for the previous transfer to complete -- the wrreq_idle wait happened
only later, in cdcncm_transmit_work. A new batch started while the previous
NTB was still in flight overwrote the in-flight buffer, so the host dropped
the corrupted NTB and TX could wedge (wrreq_idle never reposted).
Fix: acquire wrreq_idle in cdcncm_send when starting a new batch
(dgramcount == 0), before formatting; drop the now-redundant wait in
cdcncm_transmit_work (a second wait on the init-to-1 semaphore would deadlock).
2. Concurrent transmit_work. cdcncm_send runs under the recursive netdev_lock and
calls cdcncm_transmit_work() synchronously in the buffer-full branch, while a
scheduled delaywork instance runs cdcncm_transmit_work() on ETHWORK -- two
different threads. Two EP_SUBMITs of the one wrreq corrupt the IN request
queue and leave the IN buffer prepared-but-unarmed (controller idle,
wrreq_idle never reposted).
Fix: wrap cdcncm_transmit_work in netdev_lock (the synchronous caller already
holds this recursive nxrmutex; a delaywork instance blocks until the drain
releases it), and add an empty-batch guard (dgramcount == 0 -> return) so a
delaywork that runs after a synchronous flush emptied the batch does not seal
an empty NTB and double-submit the in-flight wrreq.
Validated on RP2350 (Pico 2 W) with CONFIG_NET_TCP_WRITE_BUFFERS=y as part of the
complete fix set: 144 dense/concurrent HTTP downloads, zero wedges, ~486 KB/s
(previously transmit hung within a few requests). On RP2350 full stability under
maximal TX density additionally requires a memory barrier between the BUFF_STATUS
clear and the AVAILABLE re-arm in the Cortex-M33 USB device driver (a separate
change); these cdcncm defects are real and the fixes correct independent of it.
Signed-off-by: Ricard Rosson <ricard@groundbits.com>
Assisted-by: Claude (Anthropic Claude Code)
Signed-off-by: Ricard Rosson <ricard@groundbits.com>
qemu-rv's S-mode/non-SMP setintstack unconditionally reloaded sp to
the top of the per-cpu interrupt stack. A trap taken while already
running on that stack rewound sp back to the same address, so the
nested trap's frame overwrote the still-live outer trap's frame.
Port the bounds check already used by the canonical setintstack in
riscv_macros.S: only move sp when it is outside the interrupt stack
range.
Other vendor chip.h files have the same unconditional-reload pattern
and are left for a follow-up.
Signed-off-by: liang.huang <liang.huang@houmo.ai>
Program names containing '-' (for example, renaming hello to
hello-world via PROGNAME) previously generated an invalid identifier
<NAME>_main when constructing the main= compiler definition and the
APP_MAIN target property used during builtin list generation. This
caused the CMake build to fail because '-' is not a valid character in
a C identifier.
This mirrors the Make-based fix (Application.mk's PROGSYM) for the
traditional build.
Introduce NAME_SYM, a sanitized copy of NAME with '-' replaced by '_',
and use it only where an internal C identifier is required: the
main= COMPILE_DEFINITIONS property and the APP_MAIN target property.
Leave NAME unchanged everywhere else, including CMake target/output
names and the APP_NAME property, where hyphens are valid.
The standalone/loadable executable path (MODULE/DYNLIB/kernel build)
does not rename main() and therefore requires no sanitization because
each executable is linked independently rather than merged into a
shared builtin image.
Testing (WSL2 Ubuntu, x86_64):
- BOARD_CONFIG=sim/nsh, CONFIG_EXAMPLES_HELLO_PROGNAME="hello-world":
clean CMake configure/build; 'hello-world' runs and prints
'Hello, World!!'
- Reverted to CONFIG_EXAMPLES_HELLO_PROGNAME="hello": reconfigured
and rebuilt; 'hello' runs and prints 'Hello, World!!' (no
regression)
Fixes#19447
Signed-off-by: Ansh Rai <anshrai331@gmail.com>
Three defects that together prevented macOS from ever mounting a
composite USBMSC function (Linux was mostly unaffected because its
probe sequence and recovery timing never exercised these paths):
1. usbmsc_setup() compared the class-request wIndex against the
compile-time constant USBMSC_INTERFACEID (= CONFIG_USBMSC_IFNOBASE,
i.e. 0) instead of the composite-assigned priv->devinfo.ifnobase.
In composite mode the MSC interface number is nonzero, so
GET MAX LUN, Bulk-Only Mass Storage Reset, and GET/SET INTERFACE
all failed the index check and stalled EP0. Standalone MSC is
unaffected (ifnobase == 0), which is why this went unnoticed.
2. usbmsc_deferredresponse() has its entire body inside
#ifndef CONFIG_USBMSC_COMPOSITE, so the deferred EP0 status stage
for MSRESET/SETINTERFACE was never sent in composite mode and the
host's Bulk-Only reset timed out. (Unreachable before fix 1 --
MSRESET used to stall at the wrong-interface check.) Compile the
body in composite mode too, but suppress the worker's deferred
response for SETCONFIGURATION there: the composite driver answers
that request itself, and a duplicate zero-length packet corrupts
the EP0 state.
3. usbmsc_cmdfinishstate() stalled the bulk IN endpoint whenever a
device-to-host command left a residue, even when the response had
already been sent and terminated by a short packet (or ZLP). The
stall is BOT-legal (USB MSC BOT 6.7.2) but gratuitous: the short
packet already ended the data phase and the residue is reported in
dCSWDataResidue. Hosts such as macOS answer any bulk-IN halt during
device probing with a full Bulk-Only reset sequence, which costs
seconds per command or aborts the probe entirely (macOS probes
MODE SENSE(6) with allocation lengths that exceed the response;
Linux's probe does not). Only halt the endpoint when nothing
terminated the data phase.
Root-cause analysis and host traces in apache/nuttx#19435.
Validated on RP2350 silicon (Raspberry Pi Pico 2 W, composite
CDC-ACM + CDC-NCM + USBMSC): GET MAX LUN answers 1 LUN (previously
EP0 stall and a garbage LUN count on macOS), MSRESET completes 10/10
(previously ETIMEDOUT), MODE SENSE(6) alloc=0xC0 returns short data
plus a CSW with dCSWDataResidue and zero bulk-IN stalls across the
exact-length suite, and macOS now mounts the volume (together with the
companion DCD fixes).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Ricard Rosson <ricard@groundbits.com>
blksize_t is currently defined as int16_t, which overflows when a
filesystem reports a block size larger than 32767 bytes. This causes
st_blksize to become zero, leading to an integer divide-by-zero when
st_blocks is calculated in stat().
Widen blksize_t to int32_t to support larger filesystem block sizes.
Update nuttx_blksize_t in include/nuttx/fs/hostfs.h to keep it
consistent with include/sys/types.h.
struct geometry.geo_sectorsize (include/nuttx/fs/ioctl.h) is also
typed blksize_t, so every debug print of that field using a 16-bit
format specifier is updated to PRId32 to match the new width:
drivers/misc/ramdisk.c, drivers/mmcsd/mmcsd_spi.c, drivers/mtd/ftl.c,
fs/driver/fs_blockmerge.c, drivers/mtd/smart.c,
drivers/usbhost/usbhost_storage.c, drivers/mmcsd/mmcsd_sdio.c,
arch/arm/src/s32k1xx/s32k1xx_eeeprom.c,
arch/arm/src/lc823450/lc823450_mmcl.c.
Signed-off-by: Ansh Rai <anshrai331@gmail.com>
Signed-off-by: root <root@LAPTOP-9C7LKDC5.localdomain>
Add the POSIX d_ino (file serial number) member to struct dirent and
populate it on every readdir() path, so portable callers (e.g. scp in
dropbear) that read dp->d_ino observe a meaningful, non-zero inode
number:
- include/dirent.h: declare d_ino in struct dirent and drop the
outdated comment claiming the field is unimplemented.
- include/nuttx/fs/hostfs.h: add d_ino to struct nuttx_dirent_s so
the hostfs ABI can carry the inode number across the VFS boundary.
- arch/sim/src/sim/posix/sim_hostfs.c: forward the host's
ent->d_ino into entry->d_ino.
- fs/vfs/fs_dir.c (read_pseudodir): copy the in-memory inode's
i_ino into entry->d_ino for the pseudo filesystem.
- fs/yaffs/yaffs_vfs.c: forward yaffs's dirent->d_ino into
entry->d_ino.
- fs/rpmsgfs: extend struct rpmsgfs_readdir_s with an 'ino' field
and propagate it across the RPC in both rpmsgfs_server (fills it
from the underlying entry) and rpmsgfs_client (writes it back to
the caller's nuttx_dirent_s).
Signed-off-by: Xiang Xiao <xiaoxiang@xiaomi.com>
A 16-bit ino_t can only address 65536 distinct file serial numbers,
which is not enough for filesystems with large directory trees and
breaks portable software (e.g. dropbear's scp) that expects a wider
inode number space. Widen ino_t (and nuttx_ino_t in the hostfs ABI)
to uint32_t to match common POSIX practice.
Update fs/rpmsgfs/rpmsgfs.h accordingly: promote the 'ino' field in
struct rpmsgfs_stat_priv_s from uint16_t to uint32_t and move 'nlink'
into the trailing 16-bit slot previously occupied by the reserved
field, keeping the overall packed-struct layout/size unchanged.
Signed-off-by: Xiang Xiao <xiaoxiang@xiaomi.com>
Move uid_t/gid_t out of the CONFIG_SMALL_MEMORY #ifdef so they are
always defined as unsigned int regardless of SMALL_MEMORY.
Update include/nuttx/fs/hostfs.h to match: drop the int16_t variants
of nuttx_gid_t/nuttx_uid_t and keep a single unsigned int definition
so the hostfs RPC ABI stays in sync with sys/types.h.
Signed-off-by: Xiang Xiao <xiaoxiang@xiaomi.com>
sys_callN() wraps a bare ecall and calls no other function, so the
compiler treats it as a leaf function: with frame pointers enabled,
it only needs to spill the caller's s0, which it places in what
up_backtrace()'s fp-chain walk assumes is ra's stack slot, while the
real ra slot is never written. sched_backtrace() then misreads that
slot as the return address for this frame, either resolving to a
bogus symbol or, if the adjacent garbage happens to look
out-of-range, terminating the backtrace early.
Add "ra" to the ecall clobber list so the compiler spills/reloads ra
around the ecall like a normal call site, keeping ra and the saved
s0 in their expected slots. Gate this on
CONFIG_FRAME_POINTER && CONFIG_SCHED_BACKTRACE, the only combination
where up_backtrace()'s fp-chain walk is both valid (FRAME_POINTER)
and actually exercised (SCHED_BACKTRACE); other configurations keep
the original "memory"-only clobber and pay no extra cost.
This only fixes the syscall boundary. Leaf functions that do not
cross a syscall (e.g. up_idle()) can still lose their ra slot the
same way and are not addressed here.
Signed-off-by: liang.huang <liang.huang@houmo.ai>