Two files each defined the same 16 byte constant, KSTACK_ALIGNMENT and
SIGTRAMP_STACK_ALIGN. Use STACKFRAME_ALIGN, which arch/xtensa/include/irq.h
already gives as 16, with the STACKFRAME_ALIGN_DOWN() of nuttx/irq.h.
STACK_ALIGNMENT is not the name to use here. It is TLS_STACK_ALIGN when
CONFIG_TLS_ALIGNED is set, which is the alignment of a thread stack and not
of a frame.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
The PMS grants and refuses physical addresses, so it never sees an access
that no MMU entry translates. The cache answered such an access with zeros
and raised nothing, and the task carried on with a value it never should
have had.
Enable EXTMEM_MMU_ENTRY_FAULT and route the Cache Invalid Access interrupt
to the handler that already serves the PMS monitors. An unprivileged task
that makes the access is terminated with SIGSEGV; a privileged one still
panics. The latch is level triggered, so it is cleared with the others.
Read the cause before the clear, so the log tells the two apart: a PMS
violation is a refused translation, an MMU entry fault is an access that was
never translated.
Give the kernel_oct configuration the addresses that examples/sandbox needs
to name its targets.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Reporting a fault can itself fault. syslog reaches memory the fault being
reported may have made unreachable, so esp32s3_pagefault_dispatch() is
re-entered from inside its own _alert() and never returns, and the console
fills with the same half-printed line forever. Found under Espressif's QEMU,
where PSRAM never initialises and the kernel build needs it; the board's
PSRAM works, so hardware does not take this path.
A fault repeating at the same address and PC is not helped by reporting it
again, so the dispatcher tries three times and then halts with interrupts
off. esp32s3_userfault_abort() clears the count through
esp32s3_pagefault_clear_repeat(): reaching it means the fault was contained,
so only unbroken recursion stops the machine, and three probes at one
address do not halt a healthy system.
Verified under QEMU: 12,958,521 bytes of output in 60 s before, four reports
and a halt after. On an ESP32-S3 DevKitC, esp32s3-devkit:kernel_oct, three
identical sandbox probes in one boot are all contained.
Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
When an unprivileged task takes a fault the system cannot recover from, it
now gets a fatal SIGSEGV and only that task ends. A fault in privileged code
still panics.
What decides it is the interrupted context, not the cause: the saved PS says
whether the fault was taken in User Mode. A list of causes would leave every
cause off the list as a way for a user task to stop the machine, and there
are many -- a divide by zero, a privileged instruction, a load/store error,
and an illegal instruction, which is how a refused fetch from kernel text
arrives on this chip (TRM v1.8 p.699: a denied external-memory access is
answered with 0xdeadbeaf instead of trapping). PS.UM is clear in a kernel
thread, in a system call made on the user's behalf and in an interrupt
handler, so those still panic. If the recoverable-fault dispatcher is
enabled it still gets first refusal on causes 28, 29 and 20, the only ones
re-executing can help.
esp32s3_userfault_abort() records the exception frame as the task's context,
dispatches SIGSEGV, and returns the redirected frame, so the vector's RFE
resumes the task in the signal trampoline, whose default action exits it.
CONFIG_ESP32S3_USERFAULT_ABORT enables it, default y wherever there is an
unprivileged world, and selects SIG_DEFAULT and SIG_SIGKILL_ACTION.
Verified on an ESP32-S3 DevKitC with a WROOM-2 module,
esp32s3-devkit:kernel_oct: a user task that writes through NULL, reads a wild
address, divides by zero, calls into a buffer of garbage or branches into
kernel text is terminated on its own, while an unrelated task keeps running.
Stack overflow is not contained. On the windowed ABI it faults inside the
window overflow handler and arrives as a double exception with PS.UM already
clear; guard pages are the answer, and separate work.
Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Separate the world split from the protected user image, give WORLD1 its own
vector table and its own PMS permissions -- including the PSRAM -- clean up
the user cache-MMU windows, and stop keeping the page pool mapped.
Folds in:
xtensa/esp32s3: separate the world split from the protected user image
xtensa/esp32s3: give the unprivileged world its own vector table
xtensa/esp32s3: give the unprivileged world its permissions
xtensa/esp32s3: clean up the user cache-MMU windows
xtensa/esp32s3: stop keeping the page pool mapped
xtensa/esp32s3: give the PSRAM its own PMS permissions
Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
In RS-485 mode the driver released the direction (DE) pin only on
TX_BRK_IDLE_DONE. That interrupt belongs to the break feature
(UART_TXD_BRK), which this driver never enables, so it never fired and
DIR stayed asserted after the first transmit. The board then kept driving
the bus and collided with every reply. The interrupt was also never
cleared, so had it fired, the handler would have re-entered forever.
TX_DONE cannot simply replace it: the upper half calls txint(false) as
soon as its software buffer drains, which disabled TX_DONE while the last
bytes were still in the FIFO (see #15888).
* Keep TX_DONE enabled in txint(false) while in RS-485 mode.
* In the handler, on TX_DONE with the software buffer and TX FIFO empty,
wait (bounded) for the transmitter FSM to go idle so the last stop bit
is not clipped, release DIR and disable TX_DONE. This is how ESP-IDF's
RS-485 half-duplex mode handles it.
* Make txempty() use FIFO count and FSM state, as ESP-IDF's
uart_ll_is_tx_idle() does. The raw TX_DONE bit reads 0 before the
first transmission and is now cleared by the handler, which would
make tcdrain() wait for its full timeout.
Tested on an ESP32-S3 board with an SP3485 transceiver (DE/RE on GPIO21)
at 115200 baud, doing Modbus RTU reads against a servo drive: without the
patch every request timed out; with it DIR drops right after the last stop
bit and all reads succeed.
Assisted-by: Claude:claude-opus-5-5
Signed-off-by: Max Kriegleder <max.kriegleder@gmail.com>
up_addrenv_fork() duplicates an address environment into freshly allocated
pages mapped at the same virtual addresses. The text, data and heap regions
of the source are walked one page at a time and copied into fresh pages hung
off the child's own directory, using the two kmap slots that
CONFIG_ARCH_KMAP_NPAGES reserves for exactly this.
xtensa_fork.c already took both paths: a child that keeps the parent's stack
addresses needs no relocation, which is what a duplicated address environment
gives it. Only the hook and the Kconfig default were missing.
fork() is offered on a kernel build, which is the only mode with per-process
address environments.
Verified on an ESP32-S3-WROOM-2 with esp32s3-devkit:kernel_oct. ostest
reports "Parent and child had independent memory" and exits with status 0.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
The page pool is carved out of the PSRAM that user processes run from, and
the external memory permissions are indexed by physical address, so a
permanent kernel window onto the pool is a window onto every process, which
no permission setting can close.
Stop mapping the pool. The kernel reaches a pool page through a small
scratch region instead, mapped for one operation and invalidated afterwards.
esp32s3_pgmap() takes a slot, esp32s3_pgunmap() releases it, and
ARCH_KMAP_VBASE and ARCH_KMAP_NPAGES describe the region. Two slots are
enough, because the deepest user is up_addrenv_fork(), which holds a source
and a destination page at once.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
The common Xtensa BUILD_KERNEL support needs the chip to say what it can do
and where its memory goes.
The chip selects the address environment options it now implements, keeps the
kernel and user heaps apart, and the linker scripts separate kernel from user
text and data so the two worlds can be given different permissions.
kernel_oct configures a board for it, with the user-program layout and the
boot ROMFS a kernel build loads its programs from. The ROMFS placeholder is
rebuilt with the image, the generated copy is ignored, and the programs are
given stack sizes and room for a fork() child.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Give the ESP32-S3 the arch_addrenv_t machinery that BUILD_KERNEL needs: a
per-process page directory built from the 64 KiB MMU pages of the chip, with
allocation, teardown, and the vaddr-to-paddr translation that the kernel uses
to reach a user buffer.
The MMU, PMS and WCL primitives are exposed as an arch API first, because the
address environment code and the protected user split both need them and
neither owns them.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Route the precise cache-attribute permission faults -- Load/Store/InstrFetch
Prohibited (EXCCAUSE 28/29/20) -- from xtensa_user() to a new dispatcher,
esp32s3_pagefault_dispatch(). On a serviced fault the register frame is
returned so the exception vector's RFE re-executes the faulting instruction;
otherwise it declines to the existing panic path. Gated by
CONFIG_ESP32S3_PAGEFAULT (default n, depends on BUILD_PROTECTED); the build
is unchanged when the option is off.
This is the recoverable-fault primitive the address-environment / demand-paging
work builds on. Proven on the ESP32-S3-DevKitC WROOM-2:
- A precise LoadProhibited carries a tracking EXCVADDR (the exact faulting
address), and RFE cleanly re-executes the faulted load on return -- verified
with CONFIG_ESP32S3_PAGEFAULT_SELFTEST (the identical instruction restarts
three times, then steps past, and the task resumes with the shell alive).
- ESP32-S3 PMS (World Controller) permission violations are NOT delivered as
these precise causes; they raise the asynchronous DRAM0/IRAM0 PMS-monitor
interrupt, so PMS is an isolation (kill) mechanism, not a restartable one.
No regression: esp32s3-devkit:knsh (WROOM-2) boots to nsh and ostest passes
with the option enabled.
Assisted-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
The CI style check runs nxstyle over every file a pull request touches,
so the files changed by the previous two commits have to comply even
where the problems were not introduced here. 504 errors in 25 files are
fixed: whitespace, blank lines, brace placement, switch/case indentation,
label indentation and comment blocks only, with no functional change.
Assisted-by: DeepSeek Harness:deepseek-flash
Signed-off-by: rongbaichuan <rongbaichuan1027@163.com>
nxsem_init(), nxsem_destroy(), nxmutex_init() and nxmutex_destroy()
always return OK, so checking the result only leaves dead code: the
compiler cannot remove it, because these are cross-translation-unit calls
and the nxrmutex_destroy() test is duplicated into every inlined call
site.
Apply the convention already established in commit a47a36bc5b (PR #7473)
to the two definitions which still test the value and to the 54 remaining
call sites. No signature or prototype is changed.
Testing: stm32f103-minimum:nsh builds with -Os without new warnings.
Assisted-by: DeepSeek Harness:deepseek-flash
Signed-off-by: rongbaichuan <rongbaichuan1027@163.com>
Arming a pin as a light-sleep wake source destroyed whatever it was
configured as, permanently.
esp_pm_gpio_wakeup_prepare() has to reconfigure each masked pin to plain
INPUT and hand it to gpio_wakeup_enable(), because the wakeup path only
supports level triggering. It then never put anything back. A pin that
was also a normal peripheral interrupt -- a sensor's data-ready line, say
-- came out of the first light sleep with its trigger mode gone and never
interrupted again. Nothing failed loudly; the device just went silent.
Fixed generically rather than per-board:
- esp_configgpio() now remembers the last attr applied to each pin, and
a new esp_getconfiggpio() hands it back. This is what lets the PM
code restore a pin without having to know what the pin is for.
- esp_pm_gpio_wakeup_prepare() saves each masked pin's attr before
overwriting it, and a new esp_pm_gpio_wakeup_restore() puts it back
as soon as esp_pm_light_sleep_start() returns.
Tied to the physical sleep/wake cycle deliberately, not to PM state
transitions. An earlier attempt used a board-level pm_register()/notify()
callback and never fired at all, because the board sits in PM_STANDBY
without transitioning back to PM_NORMAL -- there is no state change to
hang the restore on. The return from esp_pm_light_sleep_start() is the
one event that always happens exactly once per sleep.
Assisted-by: Claude:claude-opus-5
Signed-off-by: Felipe Moura <moura.fmo@gmail.com>
Add esp_get_irq() to retrieve the IRQ associated with an interrupt
handle.
This allows the ESP OS abstraction to recover the IRQ when freeing an
interrupt from its handle.
The corresponding change in esp-hal-3rdparty is required to use this
API when freeing interrupts.
Related: #20216
Signed-off-by: Ahmed Ashraf NourEldeen <a.programmer55559@gmail.com>
esp_pmstandby() fed up_step_idletime() the sleep duration it *asked* for
(time_in_us) rather than the one it actually got (rtc_diff_us), and did so
unconditionally. Both halves are wrong.
esp_pm_light_sleep_start() already stalls and restores the systimer
itself, but only where SOC_SLEEP_SYSTIMER_STALL_WORKAROUND is defined --
esp32c3 and esp32p4. On every other SoC, esp32s3 included, the systimer
keeps counting straight through light sleep, so the time is already in
the clock and stepping it again adds it twice.
Measured on an esp32s3-xiao: over 54 min with 1919 light sleeps totalling
454.7 s, the monotonic clock ran 443.2 s fast -- 0.97 of the time slept,
i.e. counted exactly twice, leaving the clock 13.8% fast. Anything that
reconstructs wall time from CLOCK_MONOTONIC inherits that error; for this
collar it corrupted every IMU sample timestamp.
Invisible until light sleep started happening for real, because a board
that never sleeps never steps the clock.
Note for upstream: the risc-v copy here only switches to the measured
duration and does not gate on SOC_SLEEP_SYSTIMER_STALL_WORKAROUND. The
two should be reconciled before this is proposed -- it is kept as-is so
the asymmetry is visible rather than silently decided.
Signed-off-by: Felipe Moura <moura.fmo@gmail.com>
Assisted-by: Claude:claude-opus-5
esp_gpio_irq() registers per-pin GPIO interrupts through
gpio_isr_handler_add(), never through esp_setup_irq(), so
esp_get_handle() never finds them and up_disable_irq()/up_enable_irq()
silently no-op for any GPIO-derived irq number. Fall back to
esp_gpioirqdisable()/esp_gpioirqenable() (translating irq back to a
pin via ESP_IRQ2PIN()) when the normal interrupt-matrix lookup misses.
This surfaced through drivers/sensors/lsm6ds3trc_uorb.c: its ISR
schedules a worker to drain the sensor's FIFO over I2C and disables
its own IRQ until the worker re-enables it, so a level-triggered
source (e.g. a PM GPIO wake source left in level mode) doesn't
refire continuously and starve every task, HPWORK included, before
the worker ever gets to run. That disable/enable only works now that
up_disable_irq()/up_enable_irq() actually do something for GPIO irqs.
Signed-off-by: Felipe Moura <moura.fmo@gmail.com>
Assisted-by: Claude:claude-sonnet-5
Light sleep gates the APB clock the I2C peripheral runs on. A transfer
in flight stops mid-message and never raises its completion interrupt, so
the caller blocks in i2c_sem_waitdone() until ESP32S3_I2CTIMEOTICKS
expires and gets -ETIMEDOUT for a bus that was working perfectly.
The caller is what causes it. Blocking in i2c_sem_waitdone() is exactly
what makes the idle task runnable, and the idle task is what decides to
sleep -- so the longer the transfer, the likelier it is to be cut in half
by its own wait. Nothing about this is driver-specific.
Seen on an esp32s3-xiao reading an LSM6DS3TR-C FIFO: 6000 bytes in one
transaction, some 135 ms of bus time at 400 kHz, failing with -110 over
and over. A WHO_AM_I probe and the FIFO status read, both short, never
failed once in the same runs -- only the long burst did.
The consequences went well past one failed read. With the FIFO left
undrained the sensor's level-triggered INT1 stayed asserted, the worker
was re-entered the moment the IRQ was re-enabled, and that hot loop
starved every other task until the board wedged with no console output
and no crash dump.
pm_stay(PM_IDLE_DOMAIN, PM_IDLE) is the lightest lock that suffices:
greedy_governor_checkstate() walks up from PM_NORMAL and stops at the
first state holding a wakelock, so a stay at PM_IDLE keeps the domain out
of PM_STANDBY and PM_SLEEP while still allowing the plain WFI idle.
There is no early return between the stay and the relax.
Validated over 3 h 45 of continuous acquisition across two sessions:
wakes and drains stayed 1:1 (302/302, then 375/375), zero I2C failures of
any kind, and light sleep itself unaffected -- 11.8% of wall time asleep
in both, median sleep 2.08 s.
Signed-off-by: Felipe Moura <moura.fmo@gmail.com>
Assisted-by: Claude:claude-opus-5
/proc/pm/state0 reported a flat 0 s in its SLEEP column on a board that
was demonstrably light-sleeping, because this port never told the PM core
it had slept.
pm_stats() (drivers/power/pm/pm_changestate.c) splits the time since the
last transition into dom->wake[state] or dom->sleep[state] depending on
whether the state it is handed is PM_RESTORE. up_idlepm() called
esp_pmstandby() and carried straight on, so every second -- including the
ones spent in light sleep -- was billed to wake[]. The statistics
CONFIG_PM_PROCFS advertises were simply never true here.
Read from an esp32s3-xiao that had just spent 89 s in PM_STANDBY:
DOMAIN0 WAKE SLEEP TOTAL
standby 89s 83% 0s 0% 89s 83%
Only PM_STANDBY needs this. PM_SLEEP is deep sleep and does not return
at all -- the chip resets -- so there is nothing to attribute on its way
back.
pm_changestate(domain, PM_RESTORE) is the documented way to say this: it
skips the driver prepare/veto phase, records the statistic, notifies
drivers of the restore, and deliberately does not overwrite the domain's
state, so the domain stays in PM_STANDBY as it should.
Signed-off-by: Felipe Moura <moura.fmo@gmail.com>
Assisted-by: Claude:claude-opus-5
up_idlepm() put the domain back in PM_NORMAL with pm_changestate() but
left its local oldstate holding whatever it was before sleeping, usually
PM_STANDBY. The pm_checkstate() below then returned PM_STANDBY again,
the "newstate != oldstate" test compared PM_STANDBY against a stale
PM_STANDBY, and the whole block was skipped -- including the
esp_pmstandby() call that is the only thing in here that ever sleeps.
So after the very first wakeup the board reported PM_NORMAL essentially
forever, and light-slept only when something else happened to perturb
oldstate, such as an application taking and releasing a PM_IDLE wakelock
around a transmission window.
Measured on the esp32s3-xiao collar before this fix: 4.1 s of actual
light sleep in 2 h of near-total idleness, a 1780:1 awake-to-asleep
ratio. After it: ~13.5% of wall time asleep, thousands of sleeps, no
storms.
The dead "newstate = PM_NORMAL" assignment that used to sit here was
presumably meant to be this; it is overwritten by pm_checkstate() a few
lines below and never had any effect.
Note that fixing this is what exposed two further bugs that had been
dormant behind a board that never slept: the systimer double-count in
esp_pmstandby(), and I2C transfers being cut in half by sleep. Both are
fixed in their own commits.
Signed-off-by: Felipe Moura <moura.fmo@gmail.com>
Assisted-by: Claude:claude-opus-5
esp_wifi_event_handler() held esp_wifi_lock() across the whole event
switch, including the esp_wlan_*_hook() calls
(WIFI_EVENT_STA_CONNECTED/_DISCONNECTED, WIFI_EVENT_AP_START/_STOP).
Those hooks reach netdev_lower_carrier_on()/_off(), which take the
per-device netdev_lock().
Every other path into esp_wifi_lock() acquires the two locks in the
opposite order -- the netdev ifdown path holds netdev_lock() around
its own call into esp_wifi_api_stop(), which calls esp_wifi_lock().
An application that disconnects Wi-Fi (wpa_driver_wext_disconnect()
immediately followed by wapi_set_ifdown()) races the resulting
WIFI_EVENT_STA_DISCONNECTED callback against its own ifdown call, and
the two lock orders wedge each other permanently.
Confirmed on real ESP32-S3 hardware (XIAO ESP32-S3,
CONFIG_ESPRESSIF_WIFI + CONFIG_PM + CONFIG_SCHED_TICKLESS): the
disconnecting task and the low-priority work-queue thread each waited
on a mutex held by the other (checked live via JTAG/GDB, not inferred
from code reading alone). Reproduced 4/4 times before this fix, 0/2
after.
Fix: esp_wifi_lock() is now taken only around the specific calls that
reach into the Wi-Fi driver API (esp_wifi_scan_event_parse(),
esp_wifi_set_ps()), never spanning a esp_wlan_*_hook() call --
netdev_lock() first (or absent), esp_wifi_lock() last, on every path.
Signed-off-by: Felipe Moura <moura.fmo@gmail.com>
Assisted-by: Claude:claude-sonnet-5
up_idlepm() (esp32s3_idle.c/esp32_idle.c/esp32s2_idle.c and the
shared risc-v esp_idle.c for esp32c3/esp32c6) has a recovery branch
that forces the domain back to PM_NORMAL when oldstate is not
PM_NORMAL and nothing is currently staying at it:
pm_stay(PM_IDLE_DOMAIN, PM_NORMAL);
pm_changestate(PM_IDLE_DOMAIN, PM_NORMAL);
newstate = PM_NORMAL;
pm_stay() here has no matching pm_relax() anywhere in any of the
four files. The first time this branch runs, the stay count for
PM_NORMAL never returns to 0, and pm_checkstate() (called
unconditionally right after this block) can never recommend
anything deeper than PM_NORMAL again for the rest of uptime -- the
idle loop keeps running, but the governor is permanently pinned at
full power, with no further light or deep sleep.
Confirmed on real ESP32-S3 hardware (XIAO ESP32-S3,
CONFIG_ESPRESSIF_WIFI + CONFIG_PM + CONFIG_SCHED_TICKLESS): reading
g_pmdomains[0] live via JTAG/GDB showed a "system" wakelock stuck at
state=PM_NORMAL, count=1, acquired a few seconds after boot (right
when Wi-Fi coming up briefly moves the domain off PM_NORMAL and this
branch then forces it back). Reproduced 4/4 times before this fix
(never a single PM_STANDBY transition or light-sleep-return log line
across a 40+ minute run), 0/4 after.
The trigger is timing-dependent (whether anything else already
holds PM_NORMAL at the moment this branch runs), which is likely why
it does not reproduce on every single boot.
Fix: release the stay right after the one pm_changestate() call it
exists to force, matching the comment already there ("Keep working
in normal stage") -- a one-shot nudge, not a standing hold.
Touching the switch statement right below the fix in all four files
exposed a pre-existing nxstyle violation (case labels indented level
with the switch's opening brace instead of one level in from it, per
NuttX style); reindented alongside since checkpatch lints the whole
file. esp32s3_idle.c also had two unrelated stray-indented lines
("Perform IDLE mode power management" / up_idlepm()) in up_idle();
fixed those too, same reason.
Signed-off-by: Felipe Moura <moura.fmo@gmail.com>
Assisted-by: Claude:claude-sonnet-5
openeth_receive() (arch/xtensa/src/common/espressif/esp_openeth.c)
tracks the next expected RX descriptor in priv->cur_rx_desc, an int
initialized to 0 exactly once, in esp_openeth_initialize(). QEMU's
esp32s3 machine models the OpenCores MAC's DMA ring pointer as
resetting to descriptor 0 every time RXEN is toggled off and back on
(openeth_disable()/openeth_enable(), called from ifdown()/ifup()), but
nothing rewinds the driver's own index to match. On the very first
bring-up both start at 0, so nothing looks wrong; from the second
ifup() onward the two permanently disagree, openeth_receive() keeps
inspecting the wrong descriptor, finds it still marked "owned by HW"
(e=1), and silently drops the notification. This breaks all inbound
traffic on the interface, not just application sockets -- ARP replies
and ICMP echo replies are RX frames too, so ping breaks identically.
Re-run the same descriptor initialization esp_openeth_initialize()
does at boot -- re-arm every RX/TX descriptor, rewind
cur_rx_desc/cur_tx_desc to 0 -- inside openeth_ifup(), under the same
critical section that already toggles RXEN.
Board-independent code, and open_eth only exists as a QEMU peripheral,
so there is no real-hardware regression risk.
Assisted-by: Claude:claude-sonnet-5
Signed-off-by: Felipe Moura <moura.fmo@gmail.com>
esp_openeth_initialize() (arch/xtensa/src/common/espressif/esp_openeth.c)
attaches the MAC interrupt with esp_setup_irq() but never calls
up_enable_irq(OPENETH_IRQ_MAC), unlike every other Espressif driver in
this tree. Left masked, openeth_isr_handler() never runs and received
frames are only picked up when the netdev work thread happens to run
for some other reason (a transmit). A guest can therefore send but
effectively not receive: ping still works because each request is
itself a transmit, while a socket blocked in recvfrom() waits on a
wake-up that never comes.
Confirmed with a GDB breakpoint counter on openeth_isr_handler():
zero hits before the fix, dozens after, under QEMU's esp32s3 machine
(the open_eth NIC it emulates). With the interrupt enabled, TCP
retransmits over a fixed test window dropped from 86 to 4.
Separately, openeth_ifdown() calls openeth_enable() right under a
comment that says "Disable TX and RX" -- it should call
openeth_disable(), which is what actually disables the two DMA
descriptor rings. Fixed alongside since it's the same function and
the same class of mistake.
Board-independent (arch/xtensa/src/common/espressif), not specific
to any one esp32s3 board; open_eth itself only exists as a QEMU
peripheral, so there's no real-hardware regression risk from either
change.
Signed-off-by: Felipe Moura <moura.fmo@gmail.com>
Assisted-by: Claude:claude-sonnet-5
Add what a kernel build needs on Xtensa: a crt0 for a user process, the
kernel stack allocation that a system call switches to, the syscall entry and
return path for an unprivileged caller, and the initial register state that
starts a user task at EL0 with its save area on the kernel stack.
On the ESP32-S3 the arch code that runs while the flash mapping is in flux
moves to IRAM, and the kernel heap is placed above the user .bss so that
up_allocate_kheap() and the user address environment do not overlap.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Remove the per-arch testset implementation from the spinlock layer.
The testset abstraction predates the unified spinlock.h API and is no
longer used now that all arches provide spin_lock_irqsave()/
spin_unlock_irqrestore() directly. Drop the per-arch *_testset.{c,S}
implementations and spinlock.h files for arm, sim, sparc, tricore,
x86_64, and xtensa, along with the CXD56_TESTSET,
CXD56_TESTSET_WITH_HWSEM, and CXD56_ATOMIC_WITH_HWSEM Kconfig options
in arch/arm/src/cxd56xx, and simplify the CXD56 semaphore pool loop
in cxd56_sph.c to a single unconditional range.
Signed-off-by: zhangyu117 <zhangyu117@xiaomi.com>
A C++ module with a static object does not link. GCC registers each such
object's destructor with __cxa_atexit(dtor, obj, &__dso_handle), and
__dso_handle comes from crtbegin, which a module does not link:
hello++3.cxx:119: undefined reference to `__dso_handle'
It is reachable today with CONFIG_PIC, where a module is linked as an
executable and the symbol has to resolve. Without it the link is
relocatable, the symbol stays undefined and nothing complains until
something makes it resolve.
-fno-use-cxa-atexit registers the destructors with atexit() instead, which
puts them in .fini_array. That is also where libelf_uninit() looks for them
when the module is unloaded, so the flag that makes the link work is also
the flag that makes the destructors run.
The option goes wherever CXXELFFLAGS is defined, which is the architecture
Toolchain.defs and the boards that reassign it. The toolchains that are not
GCC or Clang are left alone: ceva, tricore, z16 and the z80 family.
The CMake build sets it once, next to where the architecture elf.cmake is
included. A generator expression keeps it off the C compiles, because the
option is valid for C++ alone and GCC warns about it otherwise, and the
compiler id gates it so that a toolchain which is neither GCC nor Clang does
not see it. It cannot go in the toolchain file itself: CMake reads that file
again inside try_compile, in a project that has not included the NuttX
extensions, so the call is an unknown command there.
Reproduced with apps/examples/elf on mps3-an547:picostest with CONFIG_PIC
enabled: hello++3 fails to link before and links after.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
When a real UART (CONFIG_UARTx_SERIAL_CONSOLE) is selected as the
system console while CONFIG_ESP32S3_USBSERIAL is also enabled (e.g. to
keep /dev/ttyACM0 available as a secondary device alongside an
external console UART), the unconditional
#ifdef CONFIG_ESP32S3_USBSERIAL
# define CONSOLE_DEV g_uart_usbserial
#endif
block silently redefines CONSOLE_DEV, clobbering the correct earlier
definition that pointed it at the chosen UART device.
Confirmed on real hardware (Seeed XIAO ESP32-S3): with UART0 selected
as console and USBSERIAL also enabled, the board boot-looped on
RTCWDT_RTC_RST every ~8s, never reaching NSH. With this fix, NSH comes
up normally over UART0 and /dev/ttyACM0 remains available.
Signed-off-by: Felipe Moura <moura.fmo@gmail.com>
Assisted-by: Claude:claude-sonnet-5
CONFIG_ESP32S3_TICKLESS hangs forever the first time a task calls a
sleep/timeout with a fractional-second component of roughly 134ms or
more (e.g. usleep(500000)) while another task is also pending a
timeout.
Root cause: NSEC_2_CTICK() computes ((nsec) * CTICK_PER_USEC) /
NSEC_PER_USEC. `nsec` (struct timespec's tv_nsec) is a 32-bit `long`,
and CTICK_PER_USEC is 16 (the S3's systimer runs at 16MHz), so the
multiplication overflows a 32-bit signed int for any tv_nsec at or
above INT32_MAX / 16 (~134,217,728 ns). The overflowed (negative)
result then gets added into up_timer_start()'s `uint64_t cpu_ticks`,
wrapping around to a value near UINT64_MAX. tickless_setcounter()
then programs the systimer alarm that many ticks in the future --
effectively never -- so nxsched_process_timer() is never called and
the waiting task sleeps forever.
Reproduced on real esp32s3-xiao hardware: apps/testing/ostest hung
indefinitely right after starting user_main(), whose first statement
is usleep(500000). Instrumented up_timer_start() to print its inputs
and observed exactly the described overflow (cpu_ticks close to
UINT64_MAX for tv_nsec=510000000). Confirmed root cause is the
concurrent-timeout case specifically: user_main's usleep() alone
works, and ostest_main's own usleep() alone works, but the two
together (matching ostest's actual task_create() + concurrent
usleep() pattern) reproduce the hang every time.
Fix: cast to uint64_t before multiplying in all three *_2_CTICK
macros, forcing 64-bit arithmetic throughout, matching how the
CTICK_2_* (division) macros are already overflow-safe.
Validated on esp32s3-xiao: with the fix, the full ostest suite (built
with CONFIG_ESP32S3_TICKLESS=y) runs past the point it used to hang
and completes end to end.
Note: while testing, ostest's own round-robin test (rr_test) failed
near the end of the run -- the two same-priority SCHED_RR threads did
not appear to interleave under tickless. That looks like a separate,
likely more architectural issue (time-slice preemption needs its own
periodic re-arm, independent of one-shot sleep timeouts) and is not
addressed by this fix; filing separately.
Signed-off-by: Felipe Moura <moura.fmo@gmail.com>
Xtensa selected neither fork primitive, so vfork() was simply absent. This
wires it onto the two-primitive semantics.
There is no assembly entry point and none is needed. Every exception entry
already runs SPILL_ALL_WINDOWS, so the whole context of the calling thread is
in its exception frame and copying its stack copies a complete frame chain.
A flat build reaches that frame through SYS_save_context, issued inline so
that the recorded stack pointer belongs to a frame that stays alive for the
whole operation; a build with syscalls reaches it through xcp.sregs, recorded
by xtensa_swint() for the duration of the call.
The stack copy needs more than a relocated stack pointer here. A windowed
ABI stores each frame's caller stack pointer absolutely, in the base save
area below the frame, so a copy taken at a different address still names the
parent throughout and the child's first retw would underflow onto the
parent's stack. xtensa_fork_rebase() walks that chain and adds the
relocation offset to each link. The copy also starts one base save area
below the stack pointer rather than at it, because the frame the child
resumes into keeps its caller's spilled a0-a3 there.
Ported from the per-architecture work, reduced to the two primitives.
Co-authored-by: Xiang Xiao <xiaoxiang781216@gmail.com>
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Add missing blank lines after declarations and fix one bad alignment
in hostfs/rpmsgfs-related files. These are pre-existing style issues
flagged by CI's whole-file nxstyle check when our PR touches these
files. No logic change (git diff -w is blank-line-only additions).
Signed-off-by: yukangzhi <yukangzhi@xiaomi.com>
nxstyle now checks case labels against their enclosing brace, and the
DWC2-derived OTG device drivers align every case label with the switch
brace itself, so any change that touches one of these files fails
checkpatch on hundreds of pre-existing lines.
Indent the switch bodies by two columns as the standard requires and
re-wrap the lines that this pushes past the width limit. Whitespace and
comment re-flow only; no code changes.
Assisted-by: Claude:claude-fable-5
Signed-off-by: Jacob Dahl <dahl.jakejacob@gmail.com>
Every DWC2-derived USB device driver enables USBSUSP in GINTMSK but not
WKUP, and every one of them ANDs GINTSTS with GINTMSK before dispatch.
The resume handler is therefore unreachable: CLASS_SUSPEND is delivered
on suspend, CLASS_RESUME never is.
For CDC/ACM that is fatal. cdcacm_suspend() calls uart_connected(false),
after which serial.c refuses every open() and write() with -ENOTCONN,
and the cdcacm_resume() that would clear it never runs. On a Linux host
with the default USB autosuspend (power/control=auto, 2000 ms) simply
closing the tty is enough to trip it, and the port stays dead for the
rest of the boot while the device remains enumerated.
Verified on STM32H7 (ARK FMU v6X): before, one host suspend leaves the
CDC/ACM port permanently -ENOTCONN; after, ten forced suspend/resume
cycles all recover with the MAVLink stream intact. The remaining
drivers carry a line-for-line copy of the same initialisation.
Assisted-by: Claude:claude-fable-5
Signed-off-by: Jacob Dahl <dahl.jakejacob@gmail.com>
This patch get some fixes from ESP-IDF to fix a rolling issue
that happens on big resolution LCDs such as 800x480 LCDs.
Signed-off-by: Alan C. Assis <acassis@gmail.com>
Assisted-by: Claude Code
Select LIBC_ATOMIC_IRQ at the architecture level (ARM7TDMI, ARM926EJS,
ARMv6M) for chips that do not support atomic operations natively. This
covers all ARM7TDMI, ARM926EJS, and Cortex-M0 based chips automatically.
Also select LIBC_ATOMIC_IRQ for specific non-ARM architectures (AVR,
RISC-V, SPARC, Xtensa) that lack atomic instruction support.
Signed-off-by: zhangyu117 <zhangyu117@xiaomi.com>
Refine the atomic Kconfig to support multiple backends:
LIBC_ATOMIC_TOOLCHAIN (compiler builtins), LIBC_ATOMIC_ARCH (arch
instructions), and LIBC_ATOMIC_IRQ (interrupt disable). Rename
arch_atomic.c to arch_atomic_irq.c since it supports the IRQ backend.
Signed-off-by: zhangyu117 <zhangyu117@xiaomi.com>
BUILD_PROTECTED defaults ESP32S3_APP_FORMAT_LEGACY to y, so a protected build
has always needed the ESP-IDF second-stage bootloader. Nothing about the
protected layout requires it: the kernel and user images are described
entirely by ESP32S3_KERNEL_OFFSET, ESP32S3_KERNEL_IMAGE_SIZE and
ESP32S3_KERNEL_RAM_SIZE, and esp32s3_userspace() maps the user image itself.
Three obstacles stood in the way.
Those three symbols were gated on ESP32S3_APP_FORMAT_LEGACY, but
protected_memory.ld needs all of them for KIROM, KDROM, UIROM, UDROM, KDRAM
and UDRAM. Without them the region lengths underflow to 2**64-1 and the
kernel/user RAM split lands nowhere, which the hardware reports as a DRAM0
PMS monitor violation once the first user process runs. The offset becomes
0x0 for simple boot, where the image is flashed at the start of the device.
protected_memory.ld had no case for a 32 MB part, so FLASH_SIZE was
undefined there and ROM, UIROM and UDROM underflowed the same way.
flat_memory.ld has had the case all along.
kernel-space.ld defined none of the symbols simple boot needs
(_image_irom_*, _image_drom_*, _bss_*), and kept none of the early code
resident. __start() runs bootloader_init() and map_rom_segments() before any
flash mapping exists, so everything they reach has to be in RAM -- including
map_rom_segments() itself, which unmaps the MMU it is running from, and
nuttx_enter_critical(), reached from rtc_clk_init() by way of regi2c. These
mirror what esp32s3_sections.ld already does for the flat build.
Verified on an ESP32-S3-WROOM-2 (32 MB octal flash), esp32s3-devkit:knsh with
FLASH_MODE_OCT: boots to NSH and runs ostest, where it reaches the same
timedmutex abort as every other target. The legacy path is untouched.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
The Espressif Wi-Fi stack cannot work unless the esp_timer subsystem has
been initialized, but nothing in the Wi-Fi code does that: it is left to
each board's bringup to call esp_hr_timer_init() first. Any board that
does not happen to make that call dies on the first RF enable.
The dependency is not visible from the Wi-Fi sources. The path is:
board_wlan_init() -> esp_wlan_sta_initialize() -> esp_wlan_initialize()
-> esp_wifi_initialize() -> esp_wifi_api_adapter_init()
and later, when the radio is first powered up:
esp_phy_enable_wrapper() -> esp_phy_enable() (esp-hal-3rdparty,
components/esp_phy/src/phy_init.c) -> phy_track_pll_init()
(components/esp_phy/src/phy_common.c)
phy_track_pll_init() calls esp_timer_create() and
esp_timer_start_periodic() wrapped in ESP_ERROR_CHECK(). Both return
ESP_ERR_INVALID_STATE while esp_timer is uninitialized, because the HAL's
own esp_timer_init_os() startup hook is compiled out on NuttX
(#ifndef __NuttX__ in components/esp_timer/src/esp_timer.c), so the timer
task and the timer ISR only ever get created from NuttX's
esp_hr_timer_init() -> esp_timer_init().
Initialize the HR Timer at the top of esp_wifi_api_adapter_init(), where
the requirement actually originates. esp_hr_timer_init() is idempotent
(it early-returns once the subsystem is up), so boards that already call
it during bringup are unaffected. Also make ESPRESSIF_WIRELESS select
ESPRESSIF_HR_TIMER explicitly instead of inheriting it through the
deprecated ESP32{,S2,S3}_RT_TIMER symbols, so the timer adapter is
guaranteed to be built whenever the radio is.
This is deliberately limited to Xtensa. The RISC-V common-espressif tree
has the same unenforced dependency, but nothing is broken there today: its
ESPRESSIF_WIRELESS already selects both ESPRESSIF_HR_TIMER and RTC_DRIVER,
and esp_rtc.c initializes the timer. The mirror change can follow from
someone able to test it on RISC-V hardware.
This was diagnosed on an out-of-tree ESP32-S3 board whose bringup lacked
the call. The failure gives no panic output at all and looks exactly like
a CPU lockup: the system tick stops, the console dies mid-line and USB
stays enumerated but unresponsive. It was tracked down with ROM-level
ets_printf() breadcrumbs along the init path plus a high-priority thread
that busy-waits on ets_delay_us(): the breadcrumb trail ends inside
phy_track_pll_init() and never reaches the print after it, and the
busy-wait thread keeps printing while every sleep()-based thread stops
waking, showing the tick is gone. Initializing the timer ahead of Wi-Fi
init makes the same image associate to an AP, obtain a DHCP lease and
serve telnet. Validated on ESP32-S3 silicon (240 MHz, no PSRAM, 16 MiB
flash).
esp32s3-devkit:wifi builds clean with the change.
Signed-off-by: Ricard Rosson <ricard@groundbits.com>
Assisted-by: Claude Opus 5 (Claude Code)
This change fixes NuttX’s CMake support when NuttX is embedded
in another project via add_subdirectory(). CMake’s CMAKE_SOURCE_DIR
and CMAKE_BINARY_DIR refer to the outermost project, causing NuttX
to access its .config, generated files, host tools, and build artifacts
in the parent project’s directories. The fix introduces NUTTX_DIR and
NUTTX_BINARY_DIR, based on CMAKE_CURRENT_SOURCE_DIR and
CMAKE_CURRENT_BINARY_DIR, and consistently uses them for NuttX
self-references while preserving existing standalone builds. It fixes
the Kconfig initialization failure reported in #19697 and allows an
embedded sim:nsh build to configure, build, and boot successfully.
The change affects only the CMake build system (not Make or Kconfig
defaults), requires the corresponding nuttx-apps change, and does not
extend add_subdirectory() support to cross-compiled non-sim boards due
to CMake’s toolchain-file limitation.
Fixes#19697.
Assisted-by: Claude:claude-sonnet-5
Signed-off-by: Alan Carvalho de Assis <acassis@gmail.com>
The disconnect handler only reconnects when the reported reason is
WIFI_REASON_ASSOC_LEAVE, so an AP-initiated deauth (beacon timeout, auth or
assoc expire) leaves the station down forever. Restore the intent flag the
driver used before 1f7c3a32e5 and 20ff68bd65, matching the ESP-IDF rule of
reconnecting unless the disconnection was requested locally.
Signed-off-by: Felipe Moura <moura.fmo@gmail.com>
During the Toybox port to NuttX, Claude noticed that changes in the
menuconfig weren't taking affect. This issue exists for a long time on
NuttX, in fact BayLibre's presentation from 2017 make jokes about our
building system not been reliable:
https://www.youtube.com/watch?v=XUJK2htXxKw&t=320s
Stale archive members from $(AR)'s additive-only behavior can linger
after Kconfig toggles change which files provide a symbol, causing dead
weight or "multiple definition" link errors on incremental builds.
Fixed by splitting ARCHIVE into two macros: ARCHIVE keeps the original
additive behavior for apps/libapps.a, which many independent
subdirectories contribute to across a build, while the new
ARCHIVE_REBUILD deletes then archives for the far more common case
of a single Makefile building its own self-contained $(OBJS)
- all 39 such call sites now use it.
Assisted-By: Claude Sonnet 5
Signed-off-by: Alan C. Assis <acassis@gmail.com>
host_ioctl() reports unsupported ioctl requests from hostfs backends. Use
-ENOTTY for that case instead of -ENOSYS so callers can distinguish an
unsupported ioctl request from a missing host operation.
Keep the other host operation stubs returning -ENOSYS; this change is limited
to ioctl semantics. Update the ARM, ARM64, RISC-V, Xtensa and Windows sim
hostfs stubs to match that behavior.
Testing:
Host: Ubuntu 22.04 x86_64
- git diff --check
- make distclean
- ./tools/configure.sh -l -a ../nuttx-apps sim:nsh
- make -j16
- printf 'help\npoweroff\n' | timeout 20s ./nuttx
- make distclean
- ./tools/configure.sh -a ../nuttx-apps sabre-6quad:knsh
- make -j16
Assisted-by: Claude:Claude-Fable-5
Signed-off-by: Lingao Meng <menglingao@xiaomi.com>
Permissions (Part 2)
Description:
In kernel builds, any unprivileged process running on the NuttX
device can open /dev/efuse and attempt to read/write fuse content.
Reading the fuses may provide valuable information to an attacker
controlling the user process. The write operation, in extreme cases
where the fuse blocks are not locked, may brick the device.
DISCLAIMER: I tried to be strict with the settings, better to relax them
later if it's needed.
This is part of https://github.com/apache/nuttx/issues/19410
See https://github.com/apache/nuttx/issues/19410
Compiles ok.
Signed-off-by: Catalin Visinescu <catalin_visinescu@yahoo.com>
Description:
In kernel builds, any unprivileged process running on the NuttX device
can open /dev/efuse and attempt to read/write fuse content. Reading the
fuses may provide valuable information to an attacker controlling the user
process. The write operation, in extreme cases where the fuse blocks are
not locked, may brick the device.
This is part of https://github.com/apache/nuttx/issues/19410
Compiles ok.
Signed-off-by: Catalin Visinescu <catalin_visinescu@yahoo.com>
All I2S driver operation functions say in their signature description
that negative errno values are returned on failure. However, some of
these same functions had `uint32_t` return types. This would result in
incorrect comparison of the return value against signed error code
values.
Signed-off-by: Matteo Golin <matteo.golin@gmail.com>
This commit updates Espressif's common source code to ensure that
critical sections are properly handled by the common source code.
Signed-off-by: Tiago Medicci Serrano <tiago.medicci@espressif.com>
Correct build errors when CONFIG_ENABLE_ALL_SIGNALS is not defined
- sched makefiles: Move pending-signal helpers from the ENABLE_ALL_SIGNALS-only
list to the !DISABLE_ALL_SIGNALS list so signal dispatch is available in
PARTIAL builds sched: make SIG_PREALLOC_ACTIONS, SIG_ALLOC_ACTIONS and
SIG_DEFAULT depend on ENABLE_ALL_SIGNALS
- sched: fix ifdefs around pending-signal queue access and signal-mask for
PARTIAL/DISABLE modes
- arch: gate SYS_signal_handler / _return calls and SYSCALL_LOOKUP(signal)
with ENABLE_ALL_SIGNALS
Signed-off-by: Jukka Laitinen <jukka.laitinen@tii.ae>
esp_rtc_rdalarm() computed tv_sec and tv_nsec from two separate
evaluations of esp_hr_timer_time_us() + offset + deadline. The
high-resolution timer advances between the two calls, so a read that
straddles a second boundary yields an inconsistent timespec.
Compute the microsecond value once into a local variable and derive
both fields from that single snapshot.
Signed-off-by: yushuailong <yyyusl@qq.com>
The RWDT register offsets were incorrectly set to ESP32-S3 values
instead of ESP32 values. This was introduced when the code was
refactored from using the local NuttX header hardware/esp32_rtccntl.h
(which had the correct offsets) to using the HAL library headers, and
the offsets were moved inline into esp32_wdt.c with wrong values.
Correct the offsets to match the actual ESP32 register layout from
soc/rtc_cntl_reg.h (RTC_CNTL_WDTCONFIG0_REG at 0x8c, INT_ENA at
0x3c, etc). Without this fix, all RWDT operations (enable, configure
timeout, enable interrupt, acknowledge interrupt, feed, write-protect)
were targeting wrong memory addresses, rendering the RWDT completely
non-functional.
Signed-off-by: Tiago Medicci Serrano <tiago.medicci@espressif.com>