Commit graph

62235 commits

Author SHA1 Message Date
Ricard Rosson
2137f35cd3 rp2040/rp23xx: order the USB BUFF_STATUS clear before the AVAILABLE re-arm
rp2040_update_buffer_control()/rp23xx_update_buffer_control() re-arm an
endpoint buffer by setting the AVAILABLE bit in the buffer-control word,
which lives in DPSRAM.  When a buffer is re-armed from the completion
path (rp2040_usbintr_buffstat/rp23xx_usbintr_buffstat), that runs after
the handler has cleared the endpoint's bit in BUFF_STATUS, which lives in
the USB controller register block -- a different peripheral region.

The bus fabric may reorder those two stores.  If the controller observes
the AVAILABLE re-arm before the BUFF_STATUS clear lands, it can transmit
the next IN packet and latch its completion in BUFF_STATUS before the
clear arrives; the late clear then wipes that just-set completion bit.
The lost completion edge stops all further buffer interrupts for the
endpoint, so the class driver's write-complete callback never runs and
TX wedges permanently.

This is most visible on RP2350 (Cortex-M33) under dense/bursty IN traffic
such as CDC-NCM with TCP write buffers.  Add a UP_DMB() at the top of the
AVAILABLE re-arm so the preceding BUFF_STATUS clear is ordered ahead of
it.  (Non-SMP builds reduce spin_lock_irqsave to a barrier-free
up_irq_save, so nothing else orders these two stores.)

Assisted-by: Claude (Anthropic Claude Code)
Signed-off-by: Ricard Rosson <ricard@groundbits.com>
2026-08-04 01:44:41 +08:00
Ricard Rosson
9c8d824b52 rp2040/rp23xx: implement stall queueing and GET STATUS responses
Two device-controller defects that break standard host error recovery,
found while root-causing why macOS never mounts a NuttX mass-storage
function (full analysis and host traces in apache/nuttx#19435):

1. epsubmit() aborted (-EBUSY) any IN request submitted while the
   endpoint was halted.  The mass storage class halts bulk-IN for a
   failed device-to-host command (BOT-legal) and then submits the CSW;
   the CSW was dropped on the floor, so after the host's Clear-Halt the
   endpoint NAKed forever and the host timed out (macOS: 30 s, then a
   Bulk-Only-reset/device-reset spiral until it disables the port).
   This is exactly the race documented for years in the usbmsc_scsi.c
   header (David Hewson's analysis); the USBMSC_STALL_RACEWAR sleep
   workaround only survives hosts that clear the halt within 100 ms --
   Linux does, macOS does not, which is why Linux testing never saw it.

   Implement the stall-queueing contract instead: requests submitted
   while an endpoint is halted are queued without arming the hardware
   (arming rewrites the buffer control word and would silently clear
   the STALL bit); a halt terminates any in-flight IN transfer (its
   hardware buffer is disarmed by the STALL write and would never
   complete); clearing the halt resets the data toggle and re-arms the
   head of the queue; stale buffer completions latched for transfers
   aborted by a halt are ignored.  RP2040/RP23XX now select
   ARCH_USBDEV_STALLQUEUE, which also retires the RACEWAR's two 100 ms
   sleeps per failed command.

2. The USB_REQ_GETSTATUS handler in ep0setup() validated the request
   but never queued the two-byte response, for all three recipients
   (device/interface/endpoint), so EP0 NAKed the host's data stage
   forever and every GET STATUS timed out.  macOS issues GetPipeStatus
   = GET STATUS(endpoint) on a halted pipe before running Bulk-Only
   reset recovery and hit this on every probe; lsusb -v's device-status
   query hangs on it as well.  Send the response: endpoint recipient
   reports the halt bit, device recipient reports self-powered,
   interface reports zeros; the status stage is armed by handle_zlp()
   exactly as for class-dispatched IN transfers.

Validated on RP2350 silicon (Raspberry Pi Pico 2 W, composite
CDC-ACM + CDC-NCM + USBMSC): pre-fix, a raw-usbfs replay of macOS's
sequence and timing reproduced both defects deterministically (CSW read
ETIMEDOUT after a delayed clear-halt; all GET STATUS variants
ETIMEDOUT).  Post-fix: the CSW survives the halt and is delivered after
Clear-Halt with correct tag/status/residue across a post-stall delay
sweep of 0-1000 ms; all GET STATUS variants answer immediately with
correct halt reporting; no regressions in the exact-length SCSI suite,
Bulk-Only reset, FAT mount and full reads, CDC-NCM/ACM, or warm
reboots; and macOS now mounts the volume (together with the companion
usbmsc fixes).  The rp2040 driver shares the code and receives the
identical fix (build-tested).

Assisted-by: Claude (Anthropic Claude Code)
Signed-off-by: Ricard Rosson <ricard@groundbits.com>
2026-08-04 01:44:41 +08:00
Ricard Rosson
8039e006d6 rp2040/rp23xx: preserve SIE_CTRL.PULLUP_EN in usbdev_register
usbdev_register() calls CLASS_BIND, which ends with DEV_CONNECT
(composite_bind/cdcacm_bind) and sets SIE_CTRL.PULLUP_EN, and then
performs a wholesale putreg32 of SIE_CTRL to set EP0_INT_1BUF --
clobbering the pull-up microseconds after it was asserted.

Enumeration only ever succeeded because the host happened to latch the
microsecond pull-up blip and issued a bus reset, whose handler
(CLASS_DISCONNECT -> DEV_CONNECT) re-arms the pull-up.  A warm host
port catches the blip; a cold-plugged port is still in attach debounce,
misses it, and never resets -- PULLUP_EN stays 0 forever and the device
is totally silent on the bus while NuttX runs normally underneath.
This presented as an intermittent, image-dependent "cold boot brick"
(boot timing shifts the blip in or out of the host's blind window).

Fix: set EP0_INT_1BUF with setbits_reg32 so PULLUP_EN survives.

Validated on RP2350 silicon (Raspberry Pi Pico 2 W): an image that
failed 0/10 cold plugs enumerated 10/10 with the fix; a second board
that had never enumerated at all was recovered by it.  The rp2040
driver has the identical code and receives the identical fix
(build-tested; the RP2350 validation exercised the shared logic).

Fixes apache/nuttx#19434

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Ricard Rosson <ricard@groundbits.com>
2026-08-04 01:44:41 +08:00
Ricard Rosson
f338a53222 arch/rp23xx: apply the same bulk/notify endpoint fixes as rp2040
rp23xx_usbdev.c is a line-for-line copy of the rp2040 USB device driver
and shares all three endpoint-handling defects fixed in the preceding
commits:

  - bulk OUT reads armed with the full request length, overflowing the
    10-bit buffer-control LEN field for large reads (e.g. cdcncm);
  - endpoint requests resubmitted from their own completion callback
    being armed twice, corrupting the data PID;
  - the buffer AVAILABLE bit written together with length/PID instead of
    afterwards.

Port the identical fixes to the RP2350 driver.  The affected functions
are byte-identical to their rp2040 counterparts, so the changes match
verbatim.  These were validated on RP2040 hardware; RP2350 shares the
same USB controller IP and driver, but was not re-tested on silicon.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AHJRvWeBMTHwzpwjaUg4HW
Signed-off-by: Ricard Rosson <ricard@groundbits.com>
2026-08-04 01:44:41 +08:00
Ricard Rosson
7ff0b8e876 arch/rp2040: set buffer AVAILABLE bit after the rest of buffer control
Per the RP2040 datasheet section 4.1.2.5.1, when handing a buffer to the
USB controller the AVAILABLE bit must be written after the rest of the
buffer-control register (length and data PID) and after a short delay,
because buffer control crosses from the system clock domain into the USB
clock domain.  Writing everything in a single store risks the controller
acting on a stale length or PID.  rp2040_update_buffer_control() wrote
the whole word, AVAILABLE included, in one access.

Follow the sequence the datasheet (and the Pico SDK) use: write the
control word with AVAILABLE cleared, wait ~12 CPU cycles, then set
AVAILABLE.  The delay covers system clocks up to 12x the 48 MHz USB
clock.

Validated on raspberrypi-pico (RP2040) as part of bringing up cdcncm;
no regression on cdcacm/usbmsc.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AHJRvWeBMTHwzpwjaUg4HW
Signed-off-by: Ricard Rosson <ricard@groundbits.com>
2026-08-04 01:44:41 +08:00
Ricard Rosson
53b5048b1f arch/rp2040: don't re-arm an endpoint request resubmitted from its callback
rp2040_txcomplete() and rp2040_rxcomplete() unconditionally called
rp2040_wrrequest()/rp2040_rdrequest() to start the next transfer after
invoking a request's completion callback.  When that callback resubmits
a request on the same, now-idle endpoint -- which cdcncm and rndis do
from their interrupt/notify completion handlers -- rp2040_epsubmit()
already arms the hardware buffer for it.  The unconditional re-arm in the
completion path then arms the same buffer a second time, toggling the
DATA0/DATA1 PID twice.  The host sees a stale PID and silently discards
the packet as a retransmission, so e.g. the cdcncm NETWORK_CONNECTION /
SPEED_CHANGE notifications never reach the host and the interface stays
NO-CARRIER.

Track whether a request's hardware buffer has already been armed with a
per-request flag (set in rp2040_wrrequest/rp2040_rdrequest, cleared in
rp2040_epsubmit) and skip the redundant re-arm when the completion
callback has already resubmitted.  The in-progress multi-packet case
(transfer not yet complete) still continues normally.

Validated on raspberrypi-pico (RP2040): the cdcncm interrupt-IN
notification is now delivered (confirmed with usbmon) and the host
brings the link up; previously it never was.  This is also the likely
cause of the long-standing rndis control-response timeout on this
controller.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AHJRvWeBMTHwzpwjaUg4HW
Signed-off-by: Ricard Rosson <ricard@groundbits.com>
2026-08-04 01:44:41 +08:00
Ricard Rosson
1e560257e6 arch/rp2040: clamp bulk OUT read length to the endpoint max packet size
rp2040_epread() armed the DPSRAM buffer-control register with the full
usbdev request length.  That length is only correct for requests no
larger than the buffer-control LEN field, which is 10 bits wide (max
1023 bytes).  Class drivers that post larger read requests -- e.g.
cdcncm allocates a 16 KiB NTB read buffer -- overflow LEN: 16384 & 0x3ff
is 0, and the high bits corrupt the neighbouring control flags.  The
controller then sees a zero-length available buffer and completes the
transfer immediately with zero bytes, over and over, so no OUT data is
ever received (cdcncm floods "Wrong NTH SIGN, skblen 0").

The receive path already accumulates a request across multiple packets:
rp2040_rxcomplete() copies each packet, advances xfrd and re-arms via
rp2040_rdrequest() until the request is satisfied or a short packet
arrives.  So the buffer only ever needs to be armed for a single
maximum-size packet.  Clamp nbytes accordingly.  This matches the
transmit path, which already chunks to ep.maxpacket in rp2040_wrrequest.

Bulk classes with small reads (cdcacm, usbmsc) were unaffected because
their request lengths already fit in LEN, which is why the defect only
showed up on cdcncm.

Validated on raspberrypi-pico (RP2040): a CONFIG_NET_CDCNCM device that
previously received nothing now passes traffic in both directions with
0% packet loss.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AHJRvWeBMTHwzpwjaUg4HW
Signed-off-by: Ricard Rosson <ricard@groundbits.com>
2026-08-04 01:44:41 +08:00
Marco Casaroli
945d25f0c9 arch/x86: default CROSSDEV on a macOS host
arch/x86 set CROSSDEV only under Cygwin, so everywhere else the build used
the bare tool names and got the host compiler.  On Linux that is a native
gcc which can produce i486 ELF objects, which is what the board README
assumes.  On macOS it is Apple clang, and on Apple Silicon that cannot
target i386 in any form -- so there is no configuration in which the default
works and a cross toolchain is not optional.

Default it to i686-elf-, which Homebrew packages and which accepts the
-march=i486 -mtune=i486 already in ARCHCPUFLAGS.  arch/x86_64 has had
exactly this stanza for its own toolchain all along; this is the same shape.

The Cygwin assignment becomes ?= to match, so that a CROSSDEV passed in from
the environment or the command line is honoured rather than overridden.

Note for anyone tempted by the toolchain they already have: a 64-bit x86
compiler with -m32 is not a substitute unless it was built with multilib.
Homebrew's x86_64-elf-gcc compiles 32-bit objects perfectly happily and has
no 32-bit libgcc to link them against, so the entire build succeeds and only
the final link fails, on __udivdi3, __divdi3, __moddi3 and __udivmoddi4.
`x86_64-elf-gcc -print-multi-lib' prints just `.;', which is the toolchain
saying so up front.

Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>

Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 01:44:23 +08:00
Marco Casaroli
67c34e314f arch/x86_64: do not wrap the HPET oneshot on a deadline that has passed.
intel64_timer_start_absolute() computed "expected - current_us" unsigned.  A
watchdog started with a delay of zero asks for a deadline that is already
current, the subtraction wraps to nearly 2^64, and the comparator is set so
far ahead that the timer never fires.  The CONFIG_INTEL64_HPET_MIN_DELAY
clamp in intel64_oneshot_start() cannot help: the wrapped value is enormous,
not small.

Ask for the shortest delay the hardware can take instead of wrapping, and let
that existing minimum-delay logic pick it.

It presents as ostest hanging in wdog_test with no output and no fault.
apps/testing/ostest/wdog.c:281 calls wdtest_once(&test_wdog, param, 0), and
NSEC2TICK() takes the next few delays (1ns, 10ns, ...) to zero ticks as well;
wdtest_once() then spins forever in its "wait until the callback is triggered
exactly once" loop.

Impact: runtime, ARCH_INTEL64_HPET_ALARM only.

Assisted-by: Claude:claude-opus-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
2026-08-04 01:43:41 +08:00
Marco Casaroli
df0456c25f arch/x86_64: inherit the kernel low memory mapping in an address environment.
copy_kernel_mappings() copies exactly one PDPT entry, the 1GB linear window
that maps physical 0-1GB at 4GB-5GB.  The boot identity mapping of the low
4GB, which lives in PDPT entries 0-3 of g_pdpt_low (intel64_head.S:756, one
page directory per 1GB) and is where every MMIO register is reached, is not
carried over.

A kernel thread never gets an address environment of its own and
addrenv_switch() leaves the last one in place for it, so as soon as any
process exists, kernel code touching MMIO faults.  The HPET at 0xfed00000
finds it immediately -- CR2=fed000f0, in intel64_hpet_getreg() under
clock_systime_ticks() on the lpwork thread -- and any MMIO driver would.

Inherit the four boot PDPT entries.  They point at the boot page directories
rather than at copies, so anything intel64_map_region() adds later is
inherited too, and they carry no X86_PAGE_USER, so user code still cannot
reach them.

Impact: runtime, CONFIG_ARCH_ADDRENV builds only.  User-space access is
unchanged; the entries added are supervisor-only.

Assisted-by: Claude:claude-opus-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
2026-08-04 01:43:41 +08:00
Marco Casaroli
3c8cfa307d arch/x86_64: give the page allocator the physical page pool base.
mm_pginitialize() documents heap_start as "the physical address of the start
of memory region", and every x86_64 consumer of mm_pgalloc() agrees:
create_spgtables(), x86_64_get_pgtable() and up_addrenv_create() all put the
result through x86_64_pgvaddr() before touching it.  x86_64_pgvaddr() in turn
range-checks against CONFIG_ARCH_PGPOOL_PBASE (arch/x86_64/src/common/
pgalloc.h:67).  arm64's equivalent passes CONFIG_ARCH_PGPOOL_PBASE.

up_allocate_pgheap() passed CONFIG_ARCH_PGPOOL_VBASE instead, and in the
other branch X86_64_PGPOOL_BASE + X86_64_LOAD_OFFSET, which is the same
mistake spelled out.  Every page handed out was therefore a virtual address
that fell outside the pool's physical window, x86_64_pgvaddr() returned 0,
and the first x86_64_pgwipe() memset NULL.

It presents as a page fault in memset() under create_spgtables() the first
time a process address environment is created, which is loading the init
program.  qemu-intel64:knsh_romfs sets PGPOOL_PBASE=0x00c000000 and
PGPOOL_VBASE=0x10c000000, so the value passed was off by the 4GB load
offset.

Impact: runtime, CONFIG_ARCH_ADDRENV builds only (CONFIG_MM_PGALLOC).

Assisted-by: Claude:claude-opus-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
2026-08-04 01:43:41 +08:00
Marco Casaroli
5f9edeb6e6 arch/x86_64: define the missing CONFIG_ARCH_HAVE_SYSCALL.
794c325947 ("arch/x64:Syscall support is enabled by default", 2025-05-27)
changed nine guards from CONFIG_LIB_SYSCALL to CONFIG_ARCH_HAVE_SYSCALL but
never added the Kconfig symbol.  Only ARCH_HAVE_SYSCALL_HOOKS exists in tree;
ARCH_HAVE_SYSCALL itself is defined nowhere, so it is always unset and since
that commit x86_64 has had no x86_64_syscall_entry(), no x86_64_syscall(), no
IA32_LSTAR/IA32_STAR programming and no syscall stub layer in any
configuration:

  arch/x86_64/src/common/Make.defs:37     x86_64_syscall.c not compiled
  arch/x86_64/src/common/CMakeLists.txt:41           likewise
  arch/x86_64/include/irq.h:86
  arch/x86_64/src/intel64/intel64_cpu.c:247, :386
  arch/x86_64/src/intel64/intel64_head.S:83, :351, :538
  arch/x86_64/src/intel64/intel64_saveusercontext.S:108

qemu-intel64:knsh_romfs and qemu-intel64:knsh_romfs_pci are the two
CONFIG_BUILD_KERNEL configurations in tree, and neither can have worked in
that time.  They still link -- nothing references the missing pieces, so
libstubs.a is simply never pulled in -- and then die the first time user code
executes SYSCALL.

Define the symbol with the condition the code had before that commit, which
is LIB_SYSCALL:  every protected and every kernel build needs the interface,
and a flat build is left exactly as it is today.

Impact: restores the system call interface for CONFIG_BUILD_KERNEL and
CONFIG_BUILD_PROTECTED on x86_64.  CONFIG_BUILD_FLAT is unaffected --
qemu-intel64:nsh still builds with ARCH_HAVE_SYSCALL unset.

Assisted-by: Claude:claude-opus-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
2026-08-04 01:43:41 +08:00
Marco Casaroli
b1e02d5a06 arch/x86_64: do not demand a TSC frequency for a clock that has none.
g_x86_64_timer_freq is assigned only under ARCH_INTEL64_TSC_DEADLINE or
ARCH_INTEL64_TSC, and is read only by the two intel64_tsc_*.c files those
options build.  With the HPET as the system clock it stays 0, which is
correct and harmless -- but x86_64_timer_calibrate_freq() panics on 0
unconditionally, so the board dies during x86_64_lowsetup().

Require a frequency only where something needs one, which is
ARCH_INTEL64_HAVE_TSC.

The failure mode is worth recording, because it gives nothing to work from:
the PANIC() happens before x86_64_earlyserialinit(), and the panic handler
itself then triple-faults, because _assert() reads up_interrupt_context() --
a %gs-relative load -- and the GS base is not programmed until
x86_64_cpu_priv_set().  The console stays completely empty and the machine
resets.

Impact: runtime, ARCH_INTEL64_HPET_ALARM only.  Configurations with a TSC
are unchanged -- the PANIC() is still compiled for them.

Assisted-by: Claude:claude-opus-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
2026-08-04 01:43:41 +08:00
Marco Casaroli
e8f2c2eba3 arch/x86_64: make ARCH_INTEL64_HPET_ALARM buildable.
ARCH_INTEL64_HPET_ALARM is one of the three members of the "System Timer
Source" choice, but selecting it does not build.  Taking qemu-intel64:nsh on
master and moving the choice off ARCH_INTEL64_TSC_DEADLINE onto it:

  intel64/intel64_hpet_alarm.c:41:24: error:
      'CONFIG_ARCH_INTEL64_HPET_ALARM_CHAN' undeclared

ARCH_INTEL64_HPET_ALARM_CHAN lives inside "if INTEL64_HPET" and nothing
selects INTEL64_HPET.  Enabling that by hand moves the failure to link time,
because intel64_oneshot_lower.c is built only when INTEL64_ONESHOT is set:

  undefined reference to `oneshot_initialize'

Enabling INTEL64_ONESHOT as well finally reaches the real problem.
intel64_oneshot_lower.c implements the counter flavour of struct
oneshot_operations_s, which exists only with ONESHOT_COUNT:

  intel64_oneshot_lower.c: error: 'const struct oneshot_operations_s' has no
      member named 'start_absolute'
  intel64_oneshot_lower.c: error: implicit declaration of function
      'oneshot_count_init'
  intel64_oneshot_lower.c: error: initialization of
      'int (*)(struct oneshot_lowerhalf_s *, const struct timespec *)' from
      incompatible pointer type ... (four more of these)

ONESHOT, ONESHOT_COUNT and ONESHOT_FAST_DIVISION were selected by
ARCH_INTEL64_TSC_DEADLINE and by nothing else, so the other members of the
same choice could never be built.

Select the four from ARCH_INTEL64_HPET_ALARM as well.  INTEL64_ONESHOT
selects INTEL64_HPET in turn, which is what brings
ARCH_INTEL64_HPET_ALARM_CHAN into existence, so one added select closes all
three stages.

Impact: build only, and only for a configuration that could not be built
before.  No existing defconfig selects ARCH_INTEL64_HPET_ALARM.

Assisted-by: Claude:claude-opus-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
2026-08-04 01:43:41 +08:00
Michal Lenc
1e4e454357 drivers/ioexpander/iso1i813t.c: fix error startup condition
The reference manual (section 3.2.1) states the user has to
read GLERR and INTERR registers to clear their bits and release
ERR pin after the startup sequence. Error bits are set to 1 after
the startup if external power supply is used.

Not clearing the bits leads to subsequent read call errors if VBB
errors are checked.

Signed-off-by: Michal Lenc <michallenc@seznam.cz>
2026-08-04 00:36:38 +08:00
Alan Carvalho de Assis
76faae16b2 tools: fix stale archive members surviving a Kconfig-driven CSRCS change
During the Toybox port to NuttX, Claude noticed that changes in the
menuconfig weren't taking affect. This issue exists for a long time on
NuttX, in fact BayLibre's presentation from 2017 make jokes about our
building system not been reliable:
https://www.youtube.com/watch?v=XUJK2htXxKw&t=320s

Stale archive members from $(AR)'s additive-only behavior can linger
after Kconfig toggles change which files provide a symbol, causing dead
weight or "multiple definition" link errors on incremental builds.
Fixed by splitting ARCHIVE into two macros: ARCHIVE keeps the original
additive behavior for apps/libapps.a, which many independent
subdirectories contribute to across a build, while the new
ARCHIVE_REBUILD deletes then archives for the far more common case
of a single Makefile building its own self-contained $(OBJS)
- all 39 such call sites now use it.

Assisted-By: Claude Sonnet 5
Signed-off-by: Alan C. Assis <acassis@gmail.com>
2026-08-04 00:36:32 +08:00
Marco Casaroli
d4f523e9f3 drivers/syslog: fix syslog_write() returning -EIO on every write
syslog_write_foreach() compares an unsigned count against a signed
accumulator:

  size_t  nwritten     = 0;
  ssize_t nwritten_max = -EIO;
  ...
  if (nwritten > nwritten_max)
    {
      nwritten_max = nwritten;
    }
  return nwritten_max;

The usual arithmetic conversions promote nwritten_max to size_t, so -EIO
becomes 4294967291 on a 32-bit target, and the comparison is never true.
nwritten_max keeps its initial value and the function returns -EIO no
matter how many bytes actually went out.  Observed under gdb on a running
target: nwritten == 64, nwritten_max == -5, (nwritten > nwritten_max) == 0.

Most callers discard the result -- syslog() itself returns void -- so this
is normally invisible.  It becomes fatal when /dev/console is backed by
syslog_console_write(), because then stdio acts on it.
lib_fflush_unlocked() sees a negative return, sets __FS_FLAG_ERROR and
returns early, before resetting fs_bufpos.  The bytes have already been
emitted, but the buffer is never cleared, so every subsequent stdio call
re-flushes the same CONFIG_STDIO_BUFFER_SIZE bytes.  The console fills
with one repeated fragment and the system makes no further progress.

Reaching that state needs CONFIG_CONSOLE_SYSLOG=y together with no driver
claiming /dev/console ahead of syslog_console_init().  Three in-tree
defconfigs qualify: x86/qemu-i486:ostest, renesas/skp16c26:ostest and
x86_64/qemu-intel64:earlyfb.  The other 56 CONSOLE_SYSLOG configurations
have a serial console that registers /dev/console first, which is why this
has gone unnoticed.

Introduced by 1685e8ff7b ("syslog: avoid an infinite loop if one channel
fails"), which changed nwritten_max from size_t to ssize_t = -EIO so that
an all-channels-failed case could be reported.  Give nwritten the same type
so the comparison is signed, which preserves that intent: nwritten_max
stays -EIO only when no channel wrote anything.  nwritten is never negative,
so the remaining comparisons against buflen are unaffected.

Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>

Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 00:36:09 +08:00
Marco Casaroli
b903062b75 tools/mkallsyms.py: let the dependency error actually reach the user
mkallsyms.py prints "Please execute the following command to install
dependencies: pip install pyelftools cxxfilt" and then calls
os._exit(errno.EINVAL).  os._exit() terminates without flushing stdio, so
when stdout is a pipe -- which it always is under make -- the message sits in
the buffer and is discarded.  What the developer sees is:

  make[1]: *** [Makefile:65: nuttx] Error 22

with no indication of the cause anywhere in the build output.  Error 22 is
just errno.EINVAL leaking out as an exit status.  The same applies to
usage(), which exits ENOENT the same way.

Use sys.exit() in both places, which raises SystemExit and lets the
interpreter flush on the way out.  Nothing else changes:  the exit statuses
are the same, and os is no longer referenced, so the import goes too.

Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>

Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 00:36:00 +08:00
Ricard Rosson
e6c245686c net/tcp: don't accept a reset connection as connected (fixes send hang)
tcp_start_monitor() is called from accept() (net/inet/inet_sockif.c) for
each newly accepted connection.  When the peer had already closed the
connection before accept() ran, the monitor takes an early-return path so
that any read-ahead data buffered on the connection can still be drained;
it returns OK in that case.  accept() (net/socket/accept.c) then marks the
new socket _SF_CONNECTED unconditionally.

If the peer aborts the connection with an RST immediately after the
three-way handshake completes (for example any close with SO_LINGER
{1, 0}), the connection is moved to TCP_CLOSED with no buffered data, yet
accept() still hands back a socket that reports _SS_ISCONNECTED.  A
subsequent blocking send() on that socket passes the connected check,
registers a send callback and waits on its semaphore forever: the only
TCP_ABORT event was delivered before the callback existed, and no further
ACK, POLL or disconnect event is generated for a closed connection, so the
waiter is never woken.

Any server that writes before reading can hit this; the telnet daemon
(netutils/telnetd) is one example, where the accepted session task blocks
in send() and never completes.

Only return OK from the already-closed path when there is actually
read-ahead data to drain.  Otherwise the connection is dead, so fall
through to the -ENOTCONN return: accept() then fails cleanly instead of
handing back a socket wedged on a connection that will never make progress.
The graceful-close-with-pending-data case (the reason the OK path exists)
is preserved by the conn->readahead check.

Signed-off-by: Ricard Rosson <ricard@groundbits.com>
Assisted-by: Claude (Anthropic Claude Code)
2026-08-04 00:35:39 +08:00
Marco Casaroli
112b2414ac arch/arm: carry CONTROL over to the fork child on Cortex-M
In a protected build arm_svcall.c treats the caller's CONTROL as part of
the saved system call state: it stores it in xcp.syscall[].ctrlreturn on
entry and restores it from there on SYS_syscall_return.  All three
Cortex-M profiles do this -- armv6-m, armv7-m and armv8-m each define the
field in arch/arm/include/<arch>/irq.h and use it symmetrically.

arm_fork_direct() copied sysreturn and excreturn to the child but not
ctrlreturn.  The child's TCB comes from kmm_zalloc(), so the field was
zero, and CONTROL == 0 is nPRIV clear: the child returned to user space
privileged while its parent returned unprivileged.  The child ran out its
life with the MPU restrictions its parent is under silently lifted, which
is the isolation BUILD_PROTECTED exists to provide.

Nothing faults, and that is why this survived.  CONTROL == 0 also selects
MSP, which sounds like it should crash immediately, but NuttX already
runs Cortex-M threads on MSP -- the parent's saved value is 0x1, nPRIV
set and SPSEL clear -- so the two differ only in the privilege bit and
there is no stack change to trip over.  Privileged code then passes every
test unprivileged code passes, so ostest cannot see it either.

Measured on an RP2350 (Cortex-M33) in BUILD_PROTECTED, breaking at the
nxtask_start_fork() call in arm_fork_direct() during task_fork_test:
parent ctrlreturn 0x00000001, child ctrlreturn 0x00000000.  With this
change both read 0x00000001.

BUILD_FLAT is unaffected: without CONFIG_LIB_SYSCALL, nsyscalls is 0 and
the whole block is skipped.  armv7-a and armv7-r are unaffected too; they
carry cpsr instead, and that is already copied.

Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
2026-08-04 00:35:32 +08:00
yushuailong
9df4423fe1 sched/sched_critmonitor: remove duplicate preemption start block
The `to->preemp_start = current` assignment when `to->lockcount > 0`
was executed twice under CONFIG_SCHED_CRITMONITOR_MAXTIME_PREEMPTION,
once in the main preemption block and again after the csection block.
Remove the redundant duplicate.

Signed-off-by: yushuailong <yyyusl@qq.com>
2026-08-04 00:35:18 +08:00
Udit Jain
6fd604de6a libs/libbuiltin/compiler-rt: skip unsupported Arm VFP builtins asm
The compiler-rt builtins build globs every arm/*.S source, but several of
those hand-written assembly files require FPU features the target may not
have.  Upstream compiler-rt selects them conditionally; NuttX did not, so
BUILTIN_COMPILER_RT builds for single-precision-FPU Arm targets (e.g.
Cortex-M33, -mfpu=fpv5-sp-d16) failed to assemble with errors such as
"selected FPU does not support instruction -- vadd.f64".

Filter the source list to match the configured FPU, in both the Makefile
and CMake builds:
  - chkstk.S / chkstk2.S are Windows/MinGW-only stack probes, always dropped;
  - with no hardware FPU (!CONFIG_ARCH_FPU) all arm/*vfp.S are dropped;
  - with a single-precision FPU (!CONFIG_ARCH_DPFPU) the double-precision
    *df*vfp.S routines are dropped.

Reproduced against compiler-rt 17.0.1 with arm-none-eabi-gcc 14.2 using the
Cortex-M33 single-precision flags: 18 of 86 arm/*.S files failed to assemble
(17 double-precision *df*vfp.S plus chkstk.S); after the filter all remaining
68 files assemble cleanly.

Fixes: https://github.com/apache/nuttx/issues/17386
Generated-by: Claude (Anthropic)
Signed-off-by: Udit Jain <uditjainstjis@gmail.com>
2026-08-03 23:18:05 +08:00
raiden00pl
c294e54a9e arch/arm/nrf52,nrf53: don't pass HCI messages under the lock
on_hci() ran the host upcall with g_sdc_dev.lock held. On nrf53 this
deadlocks the BLE link: the upcall waits for the app core,
which cannot answer while bt_hci_send() blocks on the same lock.

nrf52 modified for consistency, deadlock is not possible there.

Signed-off-by: raiden00pl <raiden00@railab.me>
Assisted-by: Claude Code
2026-08-03 23:17:25 +08:00
Marco Casaroli
52e3c8fb64 boards/mps2-an521: Correct the swapped UART0 TX and RX interrupts.
The SSE-200 subsystem the AN521 image implements lists the UART interrupts
receive first: UART0 RX is external interrupt 32 and UART0 TX is 33, which
in NuttX numbering are 48 and 49.  The configuration had those two the wrong
way round, so the TX interrupt was dispatched to uart_cmsdk_rx_interrupt.

That handler clears only UART_INTSTATUS_RX, so the asserted TX status bit
survived the acknowledgement and the interrupt re-fired immediately.  The
board live-locked in the interrupt handler as soon as the console emitted its
first character: the NSH banner appeared, and nothing ran afterwards, console
input included.

Note that the overflow interrupt already sits at 63, external 47, which is
where the SSE-200 map puts it only if the pair below it is receive first --
the two corrected values are the ones consistent with it.  mps2-an500 already
follows the same order with RX at 16 and TX at 17.

Assisted-by: Claude Code:claude-opus-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
2026-08-03 23:17:11 +08:00
Marco Casaroli
459c4b7f60 arch/rp23xx: Fix six register/bit macro name clashes.
The rp23xx hardware headers define a register address macro for every
register, then a block of register bit definitions.  In three headers a bit
definition reuses the name of a register address macro, so the register
address is silently redefined as a bit mask:

  RP23XX_POWMAN_BADPASSWD           address 0x40100000 -> (1 << 0)
  RP23XX_POWMAN_BOD_CTRL            address 0x40100018 -> (1 << 12)
  RP23XX_POWMAN_DBG_PWRCFG          address 0x401000a4 -> (1 << 0)
  RP23XX_BUSCTRL_BUS_PRIORITY_ACK   address 0x40068004 -> (1 << 0)
  RP23XX_BUSCTRL_PERFCTR_EN         address 0x40068008 -> (1 << 0)
  RP23XX_PADS_QSPI_VOLTAGE_SELECT   address 0x40040000 -> (1 << 0)

None of these headers has an in-tree user yet, which is why this has gone
unnoticed; each clash appears as a "macro redefined" warning as soon as a
driver includes the header.  Code that included one of them and used the
register by name would have dereferenced 1 or 0x1000 instead of the register.

Two of the POWMAN clashes were plain duplicates.  Per the RP2350 datasheet
BOD_CTRL bit 12 is ISOLATE and DBG_PWRCFG bit 0 is IGNORE, and the correctly
named RP23XX_POWMAN_BOD_CTRL_ISOLATE and RP23XX_POWMAN_DBG_PWRCFG_IGNORE were
already defined with the same values on the following lines, so the bare names
are simply removed.  The blank line separating the VREG_LP_EXIT and BOD_CTRL
groups is restored at the same time; its absence is what let the duplicate
hide inside the preceding group.

The other four are single field registers whose field carries no separate name
(the datasheet and the SDK describe each as a one bit register), so their bit
definitions are renamed to <REGISTER>_MASK, following the _MASK spelling these
headers already use for a field extent, and written in hex like their peers.

The rp23xx-rv copies of the three headers are identical to the arm ones and
carry the same clashes, so they get the same change and stay in sync.

No functional change: none of the six names has any user in the tree.

Assisted-by: Claude Code:claude-opus-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
2026-08-03 23:17:05 +08:00
Javier Alonso
426b9729ec Review: Address @raiden00pl comments
The fix was ported from the STM32G0 to all the STM32 platforms,
as the code is mostly the same hence presents the same failure

Signed-off-by: Javier Alonso <javieralonso@geotab.com>
2026-08-03 23:16:27 +08:00
Javier Alonso
3899064e41 stm32: Detach GPIO IRQ callbacks upon clearing "setevent"
When the interrupts get disabled, the callback(s) are still attached.
That structure is never cleared, causing several calls to attach/detach
to eventually fail as the callback queue gets full.

When the error occurs, the registration fails with error 12 (ENOMEM).
By detaching the IRQ and clearing the callbacks, this error doesn't
happen again

Signed-off-by: Javier Alonso <javieralonso@geotab.com>
2026-08-03 23:16:27 +08:00
Filipe Cavalcanti
670d39673b drivers/video: zero message on MIPI DSI driver
Use memset to clear the msg struct before using on mipi_dsi_dcs_write_buffer.

Signed-off-by: Filipe Cavalcanti <filipe.cavalcanti@espressif.com>
2026-08-03 23:16:13 +08:00
Ricard Rosson
77996d70b1 boards/rp23xx: add missing rp23xx_st7735.c LCD board glue
The rp23xx common board sources reference rp23xx_st7735.c from both
Make.defs and CMakeLists.txt under CONFIG_LCD_ST7735, but the file was
never added.  Enabling CONFIG_LCD_ST7735 on any rp23xx board therefore
fails the build on the missing source.

Add the file, modelled on the existing RP2040 sibling
boards/arm/rp2040/common/src/rp2040_st7735.c and retargeted to
rp23xx/SPI1.  It implements board_lcd_initialize(), board_lcd_getdev()
and board_lcd_uninitialize(): bring up SPI1, claim the D/C (shared with
the unused SPI1 RX pad, per the rp23xx common convention), RST and BL
pads as GPIO, pulse the panel reset and bind the ST7735 driver.

Validated on silicon on a Waveshare RP2350-LCD-0.96 (RP2350A + ST7735S
160x80 IPS): the panel powers up and displays correctly, painted at boot
with no console interaction.  Build-tested raspberrypi-pico-2:nsh with
CONFIG_LCD_ST7735=y (compiles and links).

Assisted-by: Claude (Anthropic Claude Code)
Signed-off-by: Ricard Rosson <ricard@groundbits.com>
2026-08-03 23:15:37 +08:00
Marco Casaroli
25b1fcd209 sched/addrenv: do not dereference a NULL address environment
A task does not necessarily own an address environment.  tcb->addrenv_own is
set only by addrenv_attach(), which is reached only from addrenv_allocate();
a kernel thread never allocates one, and in a protected build nothing does --
there is a single address space for the whole system and the architecture's
up_addrenv_*() are stubs.  addrenv_own is then NULL for every task, always.

That a task may have no address environment is already an expected state.
addrenv_switch() returns OK when tcb->addrenv_curr is NULL and addrenv_drop()
returns early, and every caller of addrenv_select() checks addrenv_own != NULL
before calling in:  nxsched_get_stateinfo(), nxtask_argvstr(), proc_groupenv()
and the arm, arm64, risc-v and tricore up_check_tcbstack().

addrenv_take() and addrenv_give() are the only two that dereference
unconditionally.  addrenv_join() calls addrenv_take(ptcb->addrenv_own) without
a check, so pthread_create() faults on &((struct addrenv_s *)NULL)->refs
whenever the calling task has no address environment.  With
CONFIG_DEBUG_ASSERTIONS off the same access silently corrupts low memory
instead.

Handle NULL in both, the way the rest of the file already does.
addrenv_give() returns a non-zero count for the NULL case so that callers
never conclude an absent address environment has become unreferenced and
should be destroyed.

Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 23:15:29 +08:00
Alan Carvalho de Assis
1240ee3ef7 boards/lm3s6965-ek: Disable CONFIG_READLINE_EDIT to fix build
When enabling CONFIG_READLINE_EDIT the CI fails because it reports
there is not left space in this device.

Signed-off-by: Alan C. Assis <acassis@gmail.com>
2026-08-03 23:15:19 +08:00
simbit18
b43e6ae473 boards/arm64/qemu: boards/arm64/qemu
- fix Relative file path does not match actual file.

Signed-off-by: simbit18 <simbit18@gmail.com>
2026-08-03 23:15:16 +08:00
simbit18
28feaecc6d boards/arm64/fvp-v8r: nxstyle fix Relative files path
- fix Relative file path does not match actual file.

Signed-off-by: simbit18 <simbit18@gmail.com>
2026-08-03 23:15:16 +08:00
simbit18
22442fd659 boards/arm64/am62x: nxstyle fix Relative files path
- fix Relative file path does not match actual file.

Signed-off-by: simbit18 <simbit18@gmail.com>
2026-08-03 23:15:16 +08:00
simbit18
54343a0878 boards/arm/stm32l4: nxstyle fix Relative files path
- fix Relative file path does not match actual file.

Signed-off-by: simbit18 <simbit18@gmail.com>
2026-08-03 23:15:16 +08:00
simbit18
2807492b21 boards/arm/sam34: nxstyle fix Relative files path
- fix Relative file path does not match actual file.

Signed-off-by: simbit18 <simbit18@gmail.com>
2026-08-03 23:15:16 +08:00
simbit18
f74d8e6a9d arch/tricore: nxstyle fix Relative files path
- fix Relative file path does not match actual file.

Signed-off-by: simbit18 <simbit18@gmail.com>
2026-08-03 23:15:16 +08:00
simbit18
1c6927386e arch/arm64: nxstyle fix Relative files path
- fix Relative file path does not match actual file.

Signed-off-by: simbit18 <simbit18@gmail.com>
2026-08-03 23:15:16 +08:00
Javier Alonso
4a5295a56a stm32g0: Fix FLASHIF typo
`STM32_FLASHIF_BAER` was used (which doesn't exist) instead of
`STM32_FLASHIF_BASE`

Signed-off-by: Javier Alonso <javieralonso@geotab.com>
2026-08-03 22:21:26 +08:00
hanzhijian
04ef9357e1 libc/time: Fix POSIX timezone string parsing.
When loading a zoneinfo file fails, parse a non-colon-prefixed TZ value
as a POSIX timezone string. Treat a successful tzparse() result as
success while preserving the leading-colon file-only behavior.

Assisted-by: Codex:gpt-5.6 Sol
Signed-off-by: hanzhijian <hanzhijian@zepp.com>
2026-08-03 22:21:19 +08:00
Marco Casaroli
394326b98d cmake: Do not link an executable to detect the compiler.
CMake validates a compiler by building and linking a test program.  For a
bare metal cross toolchain that link cannot succeed on its own terms, and
the flags NuttX supplies for it assume a toolchain shipped with newlib:
gcc.cmake sets

  set(CMAKE_EXE_LINKER_FLAGS_INIT "--specs=nosys.specs")

A perfectly usable arm-none-eabi-gcc built without those specs therefore
fails configuration before a single NuttX source file is considered:

  arm-none-eabi-gcc: fatal error: cannot read spec file 'nosys.specs':
  No such file or directory
  CMake Error: ... CMake will not be able to correctly generate this
  project.

Setting CMAKE_TRY_COMPILE_TARGET_TYPE to STATIC_LIBRARY makes the
detection step compile without linking, which is the documented approach
for cross compiling to a bare metal target.  Toolchains that do ship
nosys.specs are unaffected: the flag only applies to CMake's own
detection, not to the NuttX link.

This is not arch specific, so it is set once in the top level CMakeLists.txt
alongside the toolchain-file selection, before project() triggers detection.
sim is excluded: it is hosted and links real host executables, so it keeps
the usual executable-based detection.  try_compile() reads the variable from
the calling scope, so it takes effect without living in a toolchain file;
tools/toolchain.cmake.export already sets the same for the exported build.

Assisted-by: Claude Code:claude-opus-4-8
Assisted-by: Claude Code:claude-fable-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
2026-08-03 22:21:10 +08:00
Yang-Rui Li
573c32f251 arch/arm/stm32h7: invalidate dcache after aligned SDMMC RX DMA
The aligned direct-DMA receive path only invalidated the destination
buffer before the transfer in stm32_dmarecvsetup(). On the Cortex-M7
the cache can speculatively prefetch into that cacheable buffer
between the pre-DMA invalidate and DMA completion, leaving stale
lines that shadow the data just written by the IDMA, so the CPU
reads a previously cached sector instead of the freshly received
data.

Invalidate again in stm32_recvdma() once the aligned transfer
completes, before the buffer is consumed. The buffer and length are
cache-line aligned on this path, so no adjacent memory is affected.

This matches the STM32 AN4839 guidance that a cache invalidate is
required after DMA completion and before the CPU reads the updated
region, not only before the transfer starts. A related instance of
the same "invalidate too early" defect on STM32H7 SPI DMA is tracked
in apache/nuttx#11594.

Root-caused on a PX4 FMUv6C (STM32H743) board where MAVLink ULog
downloads were intermittently corrupted: forensic diffing showed
corrupted windows were exactly 32 bytes (the D-cache line size),
cache-line aligned, and byte-for-byte equal to the previous 512-byte
SD sector cached in the FAT single-sector buffer. Disabling the
D-cache made the corruption disappear, isolating the defect to cache
coherency. After this fix, downloaded files matched the source file
byte-for-byte (sha256 identical) across a 5.8 MB log spanning
thousands of sectors.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: Yang-Rui Li <yang77567789@gmail.com>
2026-08-03 22:21:02 +08:00
DuoYuWang
ec9891c03a drivers/mmcsd: fix eMMC bus width switch sequencing
Two sequencing problems in the MMC wide bus path break eMMC 4-bit
operation on hosts that program the bus width in the SDIO widebus /
clock callbacks (e.g. STM32H7):

1. The SWITCH command (CMD6) is issued before the host has switched
   to wide bus operation.  When the card completes the switch while
   the host is still in 1-bit mode the switch never takes effect and
   every following data transfer times out.  Switch the host to wide
   bus operation before issuing CMD6.

2. The transfer clock is selected only at the end of mmcsd_widebus(),
   so the whole switch sequence runs at ID-mode clock and, on the
   affected hosts, the final clock update does not take effect either,
   leaving the bus at ~400 kHz.  Select the MMC transfer clock before
   calling mmcsd_widebus(), and pick CLOCK_MMC_TRANSFER_4BIT when wide
   bus operation is active (mirroring the SD card path) so a later
   clock selection cannot revert the host to 1-bit.

No behavior change for SD cards, and no change on hosts whose widebus
callback only records the requested state.

Tested on a custom STM32H743 board with eMMC: sd_bench sequential
write ~4.1 MB/s, sequential read ~6.3 MB/s (previously all data
transfers timed out).

Signed-off-by: DuoYuWang <thirteenking.wang@gmail.com>
2026-08-03 22:20:52 +08:00
Ansh Rai
aa4facaf19 lib/math32: Avoid __uint128_t casts for LDC ImportC
Building the hello_d example with LDC ImportC fails because ImportC
does not correctly handle direct __uint128_t C-style casts such as
(__uint128_t)a and (__uint128_t)1.

Replace the direct casts with equivalent typed temporaries, preserving
the existing behavior while allowing hello_d to build successfully with
LDC ImportC.

Verified on sim:nsh with CONFIG_EXAMPLES_HELLO_D=y:

  nsh> hello_d
  Hello World, [skylake]!
  DHelloWorld.HelloWorld: Hello, World!!

Signed-off-by: Ansh Rai <anshrai331@gmail.com>
2026-08-03 22:20:43 +08:00
Luchian Mihai
4459ae2d5f arch/arm/stm32h7: fix MDIO RDA field in c22 write
Use correct register

Signed-off-by: Luchian Mihai <luchiann.mihai@gmail.com>
2026-08-03 22:20:40 +08:00
liang.huang
7667d6a9fa sched: avoid dumping raw memory at address 0 for tasks without kstack
dump_stacks() only checked kernelstack_sp != 0 || force, so for tasks
with no kernel stack (e.g. idle, kernelstack_base == 0) the force path
passed base 0 to dump_stackinfo() and dumped raw memory from address 0,
triggering a secondary fault that truncated the panic log.  Guard with
kernelstack_base != 0.

Signed-off-by: liang.huang <liang.huang@houmo.ai>
2026-08-03 21:29:04 +08:00
Junbo Zheng
311069da9c arch/arm: set PSPLIM to top of TLS region to protect it from overflow
A crash was observed when running ps: a BusFault in nxtask_argvstr
dereferencing tl_argv, because a thread's stack overflow had silently
corrupted the TLS region where tl_argv resides.

On ARMv8-M with CONFIG_ARMV8M_STACKCHECK_HARDWARE, PSPLIM was set to
stack_alloc_ptr -- the bottom of the allocation where TLS begins. The
stack grows downward and TLS occupies [stack_alloc_ptr, stack_alloc_ptr
+ tls_info_size()), so an overflow crossed into TLS and clobbered
tl_argv before SP reached the limit, going undetected until code that
read the corrupted TLS data (such as ps) hit the bad pointer.

Set the limit to stack_alloc_ptr + tls_info_size() -- the top of the
TLS region and the usable stack base -- so an overflow faults at the
TLS boundary, before any TLS byte is touched. Include <tls/tls.h>
for the tls_info_size() macro, which is the value sched reserves for the
TLS region via up_stack_frame().

Signed-off-by: Junbo Zheng <zhengjunbo1@xiaomi.com>
2026-08-03 21:28:55 +08:00
liang.huang
69d7833fba sched/backtrace: fix cross-CPU buffer access under addrenv
The remote CPU's IPI handler wrote into the caller's buffer directly,
which may not be reachable from the target CPU's address environment
under CONFIG_ARCH_ADDRENV.

Signed-off-by: liang.huang <liang.huang@houmo.ai>
2026-08-03 21:28:44 +08:00
Ansh Rai
abafe75dcf tools/Zig.defs: Fix Zig builds on sim:x86_64
Building Zig applications on sim:x86_64 with Zig 0.13.0 fails because
the generated target "x86_64-freestanding-sysv" is not recognized by
Zig. Introduce a Zig-specific ABI mapping that translates "sysv" to
"gnu" without affecting other toolchains.

After correcting the target ABI, linking still fails due to unresolved
references to __zig_probe_stack. Add -fcompiler-rt so Zig includes the
required compiler runtime during linking.

Verified on sim:nsh with CONFIG_EXAMPLES_HELLO_ZIG=y:

  nsh> hello_zig
  [sim]: Hello, Zig!

Fixes #19475

Signed-off-by: Ansh Rai <anshrai331@gmail.com>
2026-08-03 21:02:20 +08:00
Catalin Visinescu
499ff232aa drivers/eeprom: I2C EEPROM Read/Write Kernel Operations Cause a Device Crash
If user passes NULL as buffer, the driver may crash. This is problematic
for NuttX protected and kernel builds.

Details in the issue: https://github.com/apache/nuttx/issues/19473

While this fixes the issue with a typical NULL pointer, fundamentally this
will be addressed with the implementation of access_ok().

https://man7.org/linux/man-pages/man2/access.2.html

Signed-off-by: Catalin Visinescu <catalin_visinescu@yahoo.com>
2026-08-03 21:00:39 +08:00