The loader packs the allocatable sections of a fully linked program into the
text and data regions in section header order, and on a chip that selects
ARCH_HAVE_TEXT_HEAP_WORD_ALIGNED_READ every section that is not executable
goes to the data region. The template left .rodata in the text region, so
the addresses the program carries did not say where it would be loaded.
.rodata now leads the data region on such a chip, .eh_frame is placed rather
than left an orphan, and the Xtensa literal pools are gathered with the text
they belong to: a literal section is not executable, so an orphan one would
be loaded into the data region, away from the code that reads it.
The ESP32-S3 needs all three. With them the shared template lays out a user
program exactly as the board script it replaces did, section for section.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
exec() of an FDPIC module now works. The loader already places such an
object and binds it; what was missing is everything binfmt has to carry
across from the load to the running task.
The task needs the module's data base in its PIC base register. binfmt
builds a D-Space for any object with a GOT, taking the base from the .got
section address; an FDPIC object names it in DT_PLTGOT instead, which the
loader has already translated, so the two are the same idea reached by
different routes and both are what up_initial_state() installs.
Constructors are not binfmt's business. A module carries its own crt0,
which walks .init_array on the task that runs the module and then calls
main, so they run in the module's own context and with its own data base.
For a module that arrives through dlopen(), libelf_insert() walks the array
instead, and it enters each entry through fdpic_invoke() because a
descriptor resolved on the calling task carries the wrong base.
The read-only segment of a module that executes in place is held by a
filesystem pin. The load takes it, and the module owns it from the point
where nothing can fail any more; it is given back when the task that runs
the module exits. The pin is held through a reference to the file rather
than a descriptor, because the descriptor belongs to the task that called
the loader and the release happens on another one.
libelf_remove() and libelf_uninit() give back what an FDPIC module holds:
the pin, and the writable segment, while the read-only one is media rather
than an allocation and must not be freed.
Built for mps3-an547:picostest with CONFIG_FDPIC both ways.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
nxstyle wants a blank line between a declaration and the statements that
follow it. The line is not new, but it sits within three lines of the
FDPIC change in this series, so CI reads it as part of the patch.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
The base firmware and an FDPIC module disagree about what a function
pointer is. Firmware is not built FDPIC, so to it a pointer is a code
address and it branches there. A module passes the address of a two word
descriptor instead, because its code and data are placed independently and
a bare code address would leave the callee unable to find its own data. A
firmware routine that takes a callback therefore branches into the
module's data segment and faults.
So the ten entry points that can be handed a callback by a module resolve
the descriptor before storing or branching to it: qsort, bsearch,
pthread_create, signal, sigaction, task_create and task_create_with_stack,
task_spawn, pthread_once, scandir, and mq_notify and timer_create with
SIGEV_THREAD.
Which one resolves matters as much as that one does. Resolving twice would
take an already resolved code address for a descriptor and read two words
from the instruction stream, so each pointer is resolved exactly once, at
the outermost point that sees it. signal() passes its argument through
untouched because sigaction() and then nxsig_action() will resolve it,
which covers a module calling sigaction() directly as well. qsort() is
split so that the public entry resolves and the recursive implementation
does not. scandir() resolves its filter but not its comparison function,
which it hands to qsort().
Whether a caller is a module at all is asked of the PIC base register,
which up_initial_state() sets only for a task that has a D-Space. A plain
kernel task therefore reads zero and is left alone.
SIGEV_THREAD is the case the register cannot answer, because the callback
runs later on a work queue worker that carries no module's base at all.
The base is captured instead when the notification is registered, in the
module's own context, and installed around the call.
All of it is behind CONFIG_FDPIC, which defaults off. Built for
mps3-an547:picostest both ways; with it off the entry points compile to
what they were.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
nxsem_init(), nxsem_destroy(), nxmutex_init(), nxmutex_destroy(),
nxrmutex_init() and nxrmutex_destroy() cannot fail, so promising a
negated errno value on failure documents an error that is never returned.
The coding standard asks the returned value description to identify all
error values of a function, and there are none, so state that OK is
always returned.
Follows "sched/semaphore: Remove the return value check of
nxsem_init/nxmutex_init", which removed the last checks of these values.
Assisted-by: DeepSeek Harness:deepseek-flash
Signed-off-by: rongbaichuan <rongbaichuan1027@163.com>
nxsem_init(), nxsem_destroy(), nxmutex_init() and nxmutex_destroy()
always return OK, so checking the result only leaves dead code: the
compiler cannot remove it, because these are cross-translation-unit calls
and the nxrmutex_destroy() test is duplicated into every inlined call
site.
Apply the convention already established in commit a47a36bc5b (PR #7473)
to the two definitions which still test the value and to the 54 remaining
call sites. No signature or prototype is changed.
Testing: stm32f103-minimum:nsh builds with -Os without new warnings.
Assisted-by: DeepSeek Harness:deepseek-flash
Signed-off-by: rongbaichuan <rongbaichuan1027@163.com>
POSIX requires getdelim()/getline() to allocate a new buffer whenever
*lineptr is NULL, regardless of the value of *n. The previous code read
the buffer size from *n unconditionally and only fell back to the initial
size when *n was zero, so a caller that passes *lineptr == NULL together
with an uninitialized (non-zero) *n caused lib_malloc() to be invoked with
that garbage size and typically fail with ENOMEM.
Treat a NULL *lineptr the same as a zero *n: (re)allocate from the known
BUFSIZE_INIT and ignore the untrusted *n. This matches the glibc
behaviour that portable code relies on (for example toybox grep, which
calls getdelim() with an uninitialized size variable).
Signed-off-by: Xiang Xiao <xiaoxiang@xiaomi.com>
libelf_uninit() only called libelf_freesymtab() when the module had an
uninitializer. But the exported symbol table is built by
libelf_insertsymtab() for every loaded module, and nothing in the tree
sets modinfo.uninitializer anymore: modules have registered their
teardown through .fini_array since a9cb28cd23. The condition is
therefore always false and every rmmod()/dlclose() leaks the exports
array together with the strdup-ed symbol names.
Call libelf_freesymtab() unconditionally, and clear the exports
pointers next to it instead of under a vestigial procfs guard that
dates back to the removed module initializer field.
Assisted-by: OpenAI Codex
Signed-off-by: yushuailong <yyyusl@qq.com>
libelf_remove() takes the module out of the registry but never frees
the registry entry, so every successful rmmod()/dlclose() leaks
sizeof(struct module_s), the name included. The lib_free() call was
dropped by e9550783d3 when the removal path was reworked.
Free the entry after the registry lock is released.
Assisted-by: OpenAI Codex
Signed-off-by: yushuailong <yyyusl@qq.com>
Implements the pthread_sigqueue Linux extension to pthreads. Follows a
similar implementation to sigqueue, except targeting a specific thread
through nxsig_dispatch.
Signed-off-by: Matteo Golin <matteo.golin@gmail.com>
atexit_call_exitfuncs() cached its loop bound on entry
(for (idx = aehead->nfuncs - 1; idx >= 0; idx--)), while
atexit_register() appends new entries at funcs[nfuncs] and bumps nfuncs.
Any function registered by an exit handler via atexit() / on_exit() /
__cxa_atexit() lands above the cached bound and is never invoked, even
though the registration returns OK.
This contradicts the exit(3) documentation that NuttX mirrors verbatim
in its own exit() docstring (libs/libc/stdlib/lib_exit.c):
It is possible for one of these functions to use atexit(3) or
on_exit(3) to register an additional function to be executed
during exit processing; the new registration is added to the
front of the list of functions that remain to be called.
The same restructure closes a second defect: atexit_call_exitfuncs()
read and cleared the task-group-shared ta_exit list without holding
ta_lock, while atexit_register() takes it ("The following must be
atomic"). Entries are now claimed under the lock and the handler is
invoked with the lock released, so a handler re-entering
atexit_register() cannot deadlock (also safe with the non-recursive
nxmutex used here).
Evidence: exit(3) man page, DESCRIPTION -
https://man7.org/linux/man-pages/man3/exit.3.html
NuttX mirrors this passage verbatim in its own exit() docstring --
5a209a853e/libs/libc/stdlib/lib_exit.c (L65-L70)
Before:
```
A handler that registers another function during exit processing
gets a success return from atexit(), but the new function is never
invoked - it lands above the loop bound cached on entry.
```
After:
```
A registration made during exit processing runs before the older
remaining handlers (order A -> B -> C below), matching the exit(3)
guarantee, and the list is consumed under ta_lock.
```
Testing:
Simulated (sim:nsh, CONFIG_LIBC_MAX_EXITFUNS=8).
Build and run:
```
cmake -B build -DBOARD_CONFIG=sim:nsh -GNinja
cmake -S . -B build # after setting CONFIG_LIBC_MAX_EXITFUNS=8
# in build/.config (sim:nsh default is 1)
cmake --build build -j$(nproc)
(echo hello; echo poweroff) | ./build/nuttx
```
"hello" runs the test at the NSH prompt; poweroff terminates the sim.
The test was carried by apps/examples/hello/hello_main.c (scratch only,
not part of this commit); its diff:
```
--- a/examples/hello/hello_main.c
+++ b/examples/hello/hello_main.c
@@ -24,6 +24,7 @@
#include <nuttx/config.h>
#include <stdio.h>
+#include <stdlib.h>
/****************************************************************************
* Public Functions
@@ -33,8 +34,29 @@
* hello_main
****************************************************************************/
+static void handler_b(void)
+{
+ printf("ATEXIT-TEST: handler B called (registered during exit)\n");
+}
+
+static void handler_a(void)
+{
+ int ret;
+
+ printf("ATEXIT-TEST: handler A called\n");
+ ret = atexit(handler_b);
+ printf("ATEXIT-TEST: atexit(handler_b) inside A returned %d\n", ret);
+}
+
+static void handler_c(void)
+{
+ printf("ATEXIT-TEST: handler C called\n");
+}
+
int main(int argc, FAR char *argv[])
{
printf("Hello, World!!\n");
+ atexit(handler_c); /* older entry, must run LAST */
+ atexit(handler_a); /* registers handler_b during exit */
return 0;
}
```
Before the fix:
```
Hello, World!!
ATEXIT-TEST: handler A called
ATEXIT-TEST: atexit(handler_b) inside A returned 0
ATEXIT-TEST: handler C called
```
(handler B is never invoked although its registration returned 0)
After the fix:
```
Hello, World!!
ATEXIT-TEST: handler A called
ATEXIT-TEST: atexit(handler_b) inside A returned 0
ATEXIT-TEST: handler B called (registered during exit)
ATEXIT-TEST: handler C called
```
Assisted-by: Claude Code (GLM-5.3) <claude@anthropic.com>
Signed-off-by: Junbo Zheng <zhengjunbo1@xiaomi.com>
aio_suspend() checked the completion status once and then performed a
single sigtimedwait(). Any SIGPOLL delivered by an unrelated AIO
operation (one not referenced by 'list') woke the caller even though
none of the awaited requests had completed, and with a timeout the
remaining wait time was not preserved either.
Re-check the completion status after every wakeup and continue
waiting, recomputing the remaining time from the absolute deadline so
that the full timeout is honored.
Signed-off-by: wushenhui <wushenhui@xiaomi.com>
Per POSIX, aio_read() and aio_write() must return -1 and set errno to
EINVAL when the request cannot be queued (aio_reqprio < 0,
aio_offset < 0), and the error must also be retrievable via
aio_error(). Conversely, when queuing fails with a bad file
descriptor, the error belongs to the asynchronous operation: the
functions must return 0 and report EBADF through aio_error().
- Merge the offset/reqprio checks and return ERROR with errno set,
after storing the result in aio_result for aio_error().
- Drop the aio_fildes < 0 early return: a closed descriptor is now
caught by fcntl()/aio_queue() and reported through aio_result with
the function returning OK.
- aio_error(): report -EINVAL (failed validation) through errno
instead of returning it as an error value.
Signed-off-by: tengshuangshuang <tengshuangshuang@xiaomi.com>
lio_listio() never validated 'nent' against {AIO_LISTIO_MAX}, so a
batch larger than the documented limit was silently accepted, and the
hard-coded _POSIX_AIO_LISTIO_MAX value of 2 was too small for real
workloads (LTP uses 10 entries per call).
Add the FS_AIO_LISTIO_MAX Kconfig option (default 10), use it for
_POSIX_AIO_LISTIO_MAX in include/limits.h, validate 'nent' in
lio_listio(), and report the limit through sysconf(_SC_AIO_LISTIO_MAX).
Signed-off-by: tengshuangshuang <tengshuangshuang@xiaomi.com>
POSIX declares lio_listio() as:
int lio_listio(int, struct aiocb *restrict const [restrict], int,
struct sigevent *restrict);
Update the prototype in include/aio.h (and the implementation and
libc.csv entry) accordingly, and drop the parameter names from the
other aio_* prototypes for consistency.
Signed-off-by: guoshichao <guoshichao@xiaomi.com>
lio_listio() submits I/O through the internal aio_read/aio_write
helpers and is only built when CONFIG_FS_AIO is enabled. Keeping it in
libs/libc splits one subsystem across two directories and forces fs/aio
to export internal interfaces to the libc build.
Move the file (and its two build system entries) from libs/libc/aio to
fs/aio so that the whole AIO implementation lives in one place.
Signed-off-by: Xiang Xiao <xiaoxiang@xiaomi.com>
A function pointer under FDPIC is not a code address. Because each
PT_LOAD segment is placed independently, a pointer has to carry the data
base its callee will need, so it is a two-word descriptor: the entry
point, and the base to install in the PIC register before branching.
R_ARM_FUNCDESC_VALUE says "the thing you are patching is such a
descriptor", and R_ARM_FUNCDESC says "manufacture one and give me its
address".
Both need state a relocation cannot carry. A descriptor's second word is
the *object's* data base, from DT_PLTGOT, and R_ARM_FUNCDESC carves
descriptors from a pool whose cursor has to survive from one relocation
to the next. up_relocate() is handed only a relocation, a resolved
symbol and an address to patch.
arch_data is the existing channel for exactly this -- RISC-V already uses
it to remember a HI20 relocation while its LO12 partner is processed --
but nothing has ever put loader state into it: it is declared zeroed and
written only by up_relocate() itself. So ARCH_ELFDATA_INIT and
ARCH_ELFDATA_FINI are added, seeding the block from the loadinfo before
the relocation loop and reading the cursor back after. Both default to
nothing, so an architecture that does not define them is unaffected, and
RISC-V's use of arch_data is untouched. libelf_relocatedyn() walks both
dynamic tables under one arch_data, so the cursor spans the whole object.
The addend handling is the part that is easy to get wrong. REL format
keeps the addend in place, in the word about to become the entry point,
and a pointer to a static function is referenced through its *section*
symbol -- the value is the section base and the offset, including the
Thumb bit, is entirely in the addend. Dropping it yields an even address
and the core faults trying to execute it as ARM code.
The GOT written into a descriptor is the loading object's own, even for
an imported function, which is what makes a callback work: when the base
firmware's qsort() calls back into a module's comparison function, the
module needs its own data base in the PIC register.
libelf_relocatedyn()'s imported-symbol path needed a change to suit. It
stores the resolved address directly and never calls up_relocate(), which
cannot produce a two-word descriptor, so under FDPIC the resolved value
now goes through up_relocate() and the relocation type decides what to
write.
Implemented for armv7-m and armv8-m, the profiles FDPIC targets; the
other ARM variants gain the arch_data block but no new relocations.
Built and booted mps3-an547:picostest and lm3s6965-ek:qemu-nxflat, the
ELF PIC and NXFLAT users of this code, both unchanged.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
With CONFIG_FDPIC selected, a module built by apps/Application.mk is now an
FDPIC shared object. Nothing about how a module is written or built
changes: the same MODULE = m in the same Makefile, the same crt0 and the
same linker script.
Two things differ from the position independent build beside it. The
compiler is told -mfdpic -fPIC, and the link is done by an
arm-uclinuxfdpiceabi linker. The stock arm-none-eabi compiler emits correct
FDPIC objects for both C and C++, so only the link needs it: the stock
linker carries the armelf emulation alone and would turn every import into
an R_ARM_JUMP_SLOT, one word, where the ABI wants an R_ARM_FUNCDESC_VALUE,
which is two, a code address and the data base that goes with it. Such a
module links cleanly and then calls out of itself with the caller's data
base still in r9. That linker is in the CI image.
gnu-elf.ld.in gains the two segments an FDPIC module needs, under
CONFIG_FDPIC, because the loader places its read-only and writable segments
independently, and names .dynamic, because a shared object is bound through
it. The sections themselves are untouched and so are the symbols crt0.c
walks, so one script serves both and both build systems get it.
.bss moves to the end of the script, for every configuration and not only
FDPIC. It held no file content but sat ahead of .got and .dynamic, which
do, so the writable segment's p_filesz had to span it and the module file
carried the whole of .bss. A module with 16 KiB of .bss went from 26724 to
10340 bytes, and its writable segment from p_filesz 0x40ac to 0xac against
an unchanged p_memsz. The loader reads p_filesz off the media, so it read
those bytes too.
Built for mps3-an547:picostest with apps/examples/elf, CONFIG_FDPIC both
ways. With it on, every module in apps/bin is ARM FDPIC with two PT_LOAD
segments and enters at _start; hello++3, which has a static C++ object,
carries DT_INIT_ARRAY and DT_FINI_ARRAY. With it off the generated script
has no PHDRS and the modules are what they were.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
gotindex named the .got section header, and every user then reached through
shdr[] for what it actually wanted. Only one of the five wanted the index.
gotbase and gotsize say it directly. gotsize is the extent of .got and is
also what says the object has one, and gotbase is where the GOT ended up:
the placed address of .got for an ordinary object, or DT_PLTGOT for an FDPIC
one, which libelf_bind() already reads. Both are set in libelf_loadfile(),
after the sections are placed, so gotbase is the address the object will be
read at rather than the one it was linked for.
The GOT walk in libelf_loadfile() now runs only when there is a base, which
also keeps it off an FDPIC object. An FDPIC object's sections are never
placed, so .got carried a link time sh_addr there, and the walk read and
wrote through it. Its GOT is relocated through its own relocations.
The check that gates libelf_xipacquire() runs before the load, when neither
field is set, so it looks the section up by name. It hands the index it
found to libelf_loadfile(), which is the only reason that function takes
one: the object is searched once, not twice.
One behaviour changes: a .got that exists but is empty now reads as no GOT.
There is nothing for any of the five users to do with an empty one.
Built for pimoroni-pico-2-plus with CONFIG_PIC, CONFIG_ELF and
CONFIG_LIBC_ELF, and for mps3-an547:bl, which is the board that read the
index. Run on QEMU with mps3-an547:picostest, which loads PIC ELF modules
from a romfs: hello prints, and ostest reaches the timed mutex test, the
same as before the change.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
The FDPIC work touches these files, and nxstyle reports errors on the lines
around every hunk, which fails the check job. The errors are older than
this series: a switch body indented two columns too deep in elf_symbols.c,
and declarations with no blank line after them.
Whitespace and one reworded comment, no change in behaviour.
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
A module that dlopen()s a library gets back function addresses from
dlsym() and calls them. Under FDPIC a bare code address is not enough:
the callee needs its own data base as well, so what dlsym() returns has
to be a function descriptor.
The exported symbol table carries no type information -- symtab_s is a
name and a value, and its own comment says typing would have to be added
to support anything but function pointers -- so by the time dlsym() is
asked there is no way to tell a function from an object.
libelf_insertsymtab() is the last point that can: st_info is still in
hand there. So an FDPIC object's exported functions are published as the
address of a descriptor carved from the module's pool, and dlopen(),
dlsym() and the module registry need no knowledge of FDPIC at all. The
pool is sized for the dynamic symbol table as well as the relocations,
since both can draw from it.
That leaves the symbol values themselves, which were wrong for any
ET_DYN object. libelf_loadsymtab() adds the symbol's section address to
its value, which is right for ET_REL, where the section address is where
the section was actually placed and the value is relative to it. In a
shared object both are already full link-time addresses, so adding them
counts the section twice. It needs translating onto wherever the object
was placed instead.
Library data is shared between everything that dlopen()s it, because the
registry holds one instance per name. Giving each user its own copy
would mean teaching the registry about instances, which is a much larger
change to shared code; an executable loaded through exec() already gets
its own data, since that path loads a fresh copy each time.
Built and run on lm3s6965-ek with the examples/elf ROMFS; the FDPIC
module continues to load, relocate and call through its own descriptors.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Running one for the first time turned up two holes in the ET_DYN path.
Neither shows up in a build.
An undefined symbol is resolved with libelf_findglobal(), which searches
only the table of globally registered symbols. The export table that
exec() hands its caller went no further than the ET_REL path, so an
ET_DYN module could not import anything the caller supplied. Invisible
while such modules resolved everything internally; an FDPIC module
imports its libc, and every import failed with "Unable to resolve addr of
ext ref printf" although the caller had passed a table containing printf.
The export table is now threaded into libelf_relocatedyn() and consulted
when the global table has no answer, leaving the existing lookup order
intact.
A relocation naming a symbol defined inside the object was dropped
silently. The code handles a relocation with no symbol, and one against
an undefined symbol, but a defined symbol fell through both. That was
harmless while every dynamic relocation arriving here had symbol index
zero, which is the case for R_ARM_RELATIVE. FDPIC brings the first ones
that do not: a pointer to a static function is emitted against the
*section* symbol, so the value is the section base and the offset within
it -- including the Thumb bit -- is carried as the addend. Deriving a
value from the word being patched, as the no-symbol case does, would
translate that addend as though it were an address. Confirmed against a
real module: .text at 0x23c plus an addend of 0x95 gives 0x2d1, which is
the function with its Thumb bit.
Also stop libelf_symname() reporting a nameless symbol as an error. A
section symbol has no name, and libelf_findsymbol() walks the whole table
looking for optional entries such as nx_stacksize, so it meets these
routinely and checks for -ESRCH itself. At error level it printed ten or
more lines per module load and buried the diagnostics that matter.
Built and run on lm3s6965-ek with the examples/elf ROMFS. The ET_REL
test modules load as before, and an FDPIC module now loads, relocates,
resolves printf and puts from the table exec() supplied, and calls
through a function descriptor of its own.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
libelf_relocatedyn() reads the handful of DT_* tags it needs to walk the
relocation tables and ignores the rest. Three more matter now.
DT_PLTGOT is where the object's data base lives. An FDPIC module runs
with that in the PIC base register, and every function descriptor built
for it names the same base as the one its callee should run with, so
without it there is nothing to put in a descriptor's second word.
The DT_*_ARRAY tags are the constructor and destructor tables. These are
already found through the section headers a few lines further down, and
that path is kept, but the dynamic tags are the authoritative copy and an
object is not obliged to carry section headers at all. Both paths now
translate through libelf_addr(), so they agree on the answer rather than
depending on which ran last. The tag values themselves were missing from
include/elf.h and are added.
Sizing the descriptor pool has to happen here rather than later.
R_ARM_FUNCDESC asks the loader to manufacture a descriptor and hand back
its address, which means the space must exist by the time the relocation
is applied, and by then the segment has been placed. So libelf_elfsize()
reserves it behind the writable data, bounded by the relocation count --
one relocation cannot ask for more than one descriptor. That bound has
slack in it, but a descriptor is two words and modules are small, which
is cheaper than walking every relocation twice to get an exact count.
Nothing here runs for a non-FDPIC object. Built and booted
mps3-an547:picostest with no change in behaviour.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
An ET_DYN object is loaded into one allocation with its data behind its
text, because its data references sit at a fixed distance from the code
that makes them. An FDPIC object does not work that way: it reaches its
data through a base register, so the two segments can be placed wherever
suits, and the point of the format is that the read-only one is left on
the media and executed there while only the writable one is copied. One
copy of the text then serves every instance.
So libelf_load() grows a second case. The object announces itself in the
OS/ABI byte, which is noted once in libelf_loadhdrs() rather than
re-derived; e_flags cannot be used for this, as an FDPIC object's are an
unremarkable EABI version and testing them would reject every valid
module. Text is taken from the media address plus the segment's own file
offset -- the same arithmetic the ET_REL path already does with
sh_offset -- and libelf_loadfile() does not read it. If the filesystem
cannot show its media, the loader copies the text to RAM instead. The
module then loses the shared text and the flash saving, but it runs.
Obtaining that address needs two mechanisms, and they are not
interchangeable. A compacting filesystem can move a file's blocks, so it
hands out an address only with a pin that holds them still and expects
the pin back; xipfs is the one in tree. A filesystem whose layout never
changes has nothing to hold and answers FIOC_XIPBASE with a bare address;
romfs and tmpfs are those. libelf_xipacquire() asks for the pin first,
because a filesystem that needs one is not safe without it, and
libelf_unload() gives it back. The loader asks for a pin only if it can
hold one, or the pin would stay for ever.
The pin is thus not specific to FDPIC. Any module that executes in place
from a compacting filesystem takes one, and gives it back at unload.
mmap() is not used, though both filesystems implement it. The mapping
would be recorded against whichever task called the loader, while the
release happens when the module's own task exits, which is a different
group -- so the pin would outlive the module and the extent would never
become movable again.
Unloading has to change with placement: the existing path frees only
textalloc because ET_DYN had a single allocation, which would leak an
FDPIC object's data and free media the filesystem only lent us.
Nothing here runs for a non-FDPIC object; every branch is behind the flag
and the single-allocation path is untouched. Built and booted
mps3-an547:picostest, which is CONFIG_ELF with CONFIG_PIC, with no change
in behaviour.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Replace the local EINTR retry loop with nxsem_wait_uninterruptible().
This keeps the master implementation aligned with the libc semaphore API
without changing cancellation behavior.
Keep the cleanup separate so release branches where the helper is not
available to Protected user space can use the functional commit without a
downstream compatibility patch.
Assisted-by: Codex:GPT-5
Signed-off-by: DuoYuWang <thirteenking.wang@gmail.com>
Implement the handle-based create, queue, priority, cancellation, and
teardown APIs for CONFIG_LIBC_USRWORK. Custom queues use configurable
pthread worker pools while the predefined USRWORK queue remains available.
Match scheduler-backend delay, replacement, cancellation, and lifecycle
semantics. Restrict the libc backend to task context because it uses
blocking synchronization.
Tested on an STM32H7 PX4 FMUv6C with ostest wqueue in Protected user space.
Assisted-by: Codex:GPT-5
Signed-off-by: DuoYuWang <thirteenking.wang@gmail.com>
Both files put the body of the relocation switch at the same indent as the
switch braces, so nxstyle reports forty-four errors in each and any patch
whose hunks land near them fails the check job.
Giving the body its level takes the bit diagrams in the comments one column
past the line limit. The rulers say Instr rather than Instructions, which is
enough and is what the same rulers further down already do. A comment that
had no code on its line becomes a sentence of its own, and two that were a
column out are put right.
Whitespace and comments only. Compiled before and after for cortex-m7 and
cortex-m33: the disassembly is identical.
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Increase LINK_MAX from _POSIX_LINK_MAX (8) to 128 to allow
directories to have a reasonable number of subdirectories while
still enforcing a hard link limit.
Also fix pathconf(_PC_LINK_MAX) to return the actual LINK_MAX
value instead of the minimum _POSIX_LINK_MAX.
Signed-off-by: yukangzhi <yukangzhi@xiaomi.com>
CI feeds nxstyle the diff hunks with three lines of context, so style errors
that are older than this change, in the lines around the hunks, fail the
check job. They are a switch body indented two columns too deep, an
initializer brace one level in, and two declarations with no blank line
after them.
Whitespace only, no change in behaviour.
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
The ET_DYN path computes run-time addresses from link-time ones in five
places, each open-coding the arithmetic, and two of them disagree about
how: libelf_relocatedyn() adds textalloc to a relocation's r_offset in
one branch and subtracts datasec before adding datastart in the next,
while the value translation a few lines further down picks between those
two forms with an explicit test on datasec.
Collect that into libelf_addr(), which makes the test once: an address
below the data segment's link-time base belongs to text, anything at or
above it to data.
This changes nothing today. libelf_elfsize() sets
segpad = datasec - (text_vaddr + textsize)
and libelf_load() then places
datastart = textalloc + textsize + segpad
so datastart - datasec is textalloc, and the data branch reduces to
textalloc + vaddr -- exactly what the text branch returns, and exactly
what adding a single load bias did before. The two forms are the same
arithmetic written twice.
They stop being the same once text and data are placed independently,
which is what an FDPIC object requires: its two PT_LOAD segments are
relocated separately so that the read-only one can be mapped in place on
the media while only the writable one is copied. Having the translation
in one function is what makes that possible without auditing every
open-coded expression again.
Built for mps3-an547:picostest, which is CONFIG_ELF with CONFIG_PIC, and
boots identically to the same configuration without this change.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Implement atomic_lock/atomic_unlock using hwspinlock when
CONFIG_LIBC_ATOMIC_HWSPINLOCK is selected, and using up_irq_save/
up_irq_restore when CONFIG_LIBC_ATOMIC_IRQ is selected. Rename
arch_atomic_irq.c to arch_atomic.c.
The 64-bit atomic operations use spinlock (spin_lock_irqsave)
regardless of the selected backend, ensuring multi-core safety.
Signed-off-by: zhangyu117 <zhangyu117@xiaomi.com>
The link support is no longer limited to the pseudo file system and now
covers both soft (symbolic) links and hard links across the VFS. Rename
the configuration option PSEUDOFS_SOFTLINKS to the more accurate FS_LINKS
and update all references in the source, headers, Kconfig, documentation
and board defconfigs accordingly.
This is a configuration rename; any out-of-tree defconfig that still
selects PSEUDOFS_SOFTLINKS must be updated to FS_LINKS.
Signed-off-by: zhengyu16 <zhengyu16@xiaomi.com>
Replace ILLD intrinsics (__swap, __ld32, __cmpAndSwap) with inline
assembly functions (tricore_atomic_swap, tricore_atomic_cmpswap) to
remove the dependency on IfxCpu_Intrinsics.h.
Also fix the expect parameter type to use volatile void * to match
the declaration in atomic.h, avoiding type conflicts.
Signed-off-by: zhangyu117 <zhangyu117@xiaomi.com>
libelf_elfsize() takes textalign and dataalign from the section headers,
which only the ET_REL path walks. An ET_DYN object is sized from its
program headers instead, so both fields stay at zero, and the allocation
a few lines later asks for that alignment:
loadinfo->textalloc = lib_memalign(loadinfo->textalign, ...);
Zero is not a valid alignment, and every path that receives it divides by
it. mm_memalign() accepts zero as a power of two, because 0 & -0 is 0,
then takes the "alignment <= MM_ALIGN" branch and evaluates
"((uintptr_t)ptr) % alignment" in a DEBUGASSERT. With
CONFIG_MM_HEAP_MEMPOOL and a pool that fits the request the object never
reaches that branch and gets ALIGN_UP(blk, 0) instead, which is
((blk - 1) / 0) * 0.
On Cortex-M this is usually invisible: UDIV returns zero for a division
by zero unless CCR.DIV_0_TRP is set, which NuttX does not set, so the
assertion compares zero against zero and passes. It is a SIGFPE on the
simulator, and the mempool path returns a null pointer wherever the
division yields zero, which the loader reports as -ENOMEM.
Ask for a natural word when the program headers gave nothing. p_align is
the linker's page granularity, not a section requirement, so honouring it
would cost a page per module for no gain, and the sections of a shared
object need no more than a word.
Built for mps3-an547:picostest, which is CONFIG_ELF with CONFIG_PIC.
Runtime evidence on hardware follows.
Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Rename atomic_fetch_add/sub/or/and/xor to atomic_add/sub/or/and/xor
to avoid conflicts with the C/C++ standard library naming. The
atomic_fetch_xxx naming is reserved by the standard; keeping it causes
function name conflicts when source files indirectly include both
<nuttx/atomic.h> and <atomic>/<stdatomic.h>.
Signed-off-by: zhangyu117 <zhangyu117@xiaomi.com>
The reason for using builtin atomic is that in C++, when include <atomic> in <nuttx/atomic.h> easily conflicts with third-party function libraries. We wanted to completely separate the implementation of <nuttx/atomic.h>.
There are two points:
1. use builtin function directly.
2. Without the standard library implementation, need implement "atomic_fetch_xxx", leading conflicts with the standard library used by third-party programs, introducing redefinition issues and requiring name changes.
Signed-off-by: zhangyu117 <zhangyu117@xiaomi.com>
Split the atomic implementation into three files:
- arch_atomic.h: Shared macros (STORE, LOAD, etc.) using atomic_lock()/
atomic_unlock() abstraction
- arch_atomic_irq.c: 32-bit atomic functions using IRQ disable (conditional
on CONFIG_LIBC_ATOMIC_IRQ via Make.defs)
- arch_atomic64.c: 64-bit atomic functions using spinlock (always compiled,
multi-core safe). The __atomic_* functions are always provided (GCC
runtime helpers), while nx_atomic_* functions are conditional on
!CONFIG_LIBC_ATOMIC_TOOLCHAIN.
Signed-off-by: zhangyu117 <zhangyu117@xiaomi.com>
Refine the atomic Kconfig to support multiple backends:
LIBC_ATOMIC_TOOLCHAIN (compiler builtins), LIBC_ATOMIC_ARCH (arch
instructions), and LIBC_ATOMIC_IRQ (interrupt disable). Rename
arch_atomic.c to arch_atomic_irq.c since it supports the IRQ backend.
Signed-off-by: zhangyu117 <zhangyu117@xiaomi.com>
Tricore gcc does not support atomic interface but some users need to
use atomic operations, so support atomic function using tricore arch
instructions (__cmpAndSwap/__swap/__ld32).
Signed-off-by: zhangyu117 <zhangyu117@xiaomi.com>
The atomic implementation of machine/arch_atomic.c is achieved by
switching interrupts. This version does not support SMP.
Signed-off-by: zhangyu117 <zhangyu117@xiaomi.com>
The BSD string functions take a word path only when both pointers are
aligned, and a byte path otherwise. A pair at the same offset from a
boundary takes the byte path even though copying or comparing a few leading
bytes aligns both at once, since aligning one aligns the other.
Add MISALIGNED(), which asks whether two pointers disagree about where a
boundary falls, and walk an agreeing pair up to the boundary before the
existing path selection. MISALIGNED4() does the same for the 4-byte path,
so a pair that is 4-byte but not 8-byte aligned reaches the wide path
instead of the middle one. No existing line changes: the walk is a new step
ahead of the current decisions. A pair at differing offsets still takes the
byte path, since no single boundary serves both.
Measured on an EIC7700 EVB (EIC7700X, RV64GC, 1.4GHz) with the BSD string
functions selected and the RISC-V assembly ones disabled, using the
benchmark in apps#3706, medians of 3 runs in MB/s at its largest size:
equal offset aligned
memcpy 414 -> 4148 10.0x 4214 -> 4208
memcmp 41 -> 361 8.8x 362 -> 360
strncmp 28 -> 202 7.4x 207 -> 207
strcmp 42 -> 273 6.5x 278 -> 276
strncpy 377 -> 1676 4.5x 1824 -> 1748
stpncpy 376 -> 1654 4.4x 1843 -> 1724
stpcpy 551 -> 1833 3.3x 1970 -> 1939
memccpy 650 -> 2012 3.1x 2478 -> 2016
strcpy 636 -> 1837 2.9x 1678 -> 1965
Cases the walk never runs for move in both directions by up to a third, the
largest being memccpy at differing offsets, 648 -> 414. Their code is
unchanged, so that is code placement rather than an effect of the change.
The change is architecture independent but has only been measured on
RV64GC. Word size, alignment cost and byte loop codegen all differ
elsewhere, so the balance wants measuring on other architectures.
Assisted-by: Claude:claude-opus-5
Signed-off-by: Justin Hammond <justin@dynam.ac>
memcmp, strncmp and strcmp reach their word loops only when both pointers
are already on a register boundary:
or t0, a0, a1
andi t0, t0, SZREG-1
That asks more than the loops need. They load from the two pointers at
the same boundary, so what matters is that the two agree about where a
boundary falls, not that either is already on one. A pair offset by the
same amount can be walked up to the boundary a byte at a time and
compared a register at a time from there.
The union also holds far less often than the difference. For arbitrary
pointers on RV64 it is true about one time in 64 against one in eight,
and the case it rejects, two strings carved out of the same buffer, is
the common one.
Test the difference of the pointers, and walk to the boundary first.
arch_strcpy.S and arch_memcpy.S already do this. Keeping every access
aligned is not only faster here: the base ISA does not require misaligned
loads and stores to be supported at all, so a routine in a machine
directory cannot assume one will work, whatever it costs.
Measured on a 1.4 GHz rv64, source and destination misaligned by one:
before after
memcmp 32K 34.4 458.0 MB/s
strncmp 32K 32.4 253.0 MB/s
strcmp 32K 41.0 280.0 MB/s
Each of those was the rate of the byte loop the word loop was meant to
replace. Pointers that genuinely disagree still take the byte loop, and
the aligned rates are unchanged.
The measurements come from the benchmark in apache/nuttx-apps#3706.
Assisted-by: Claude:claude-opus-5
Signed-off-by: Justin Hammond <justin@dynam.ac>
The word loop walks src to a register boundary and then stores a whole
register at a time to dst, but nothing establishes that dst is on a
boundary too. Where the two pointers disagree about where a boundary
falls, every store in that loop is misaligned.
The base ISA does not require misaligned stores to be supported. Where
firmware emulates them each store traps into machine mode, and where
nothing emulates them the store faults, so this is not only a question of
speed. Measured on a 1.4 GHz rv64 that emulates them, with a 32 KB
string whose src and dst are misaligned by different amounts:
generic C 410.4 MB/s
this file 7.5 MB/s
which is around 178 cycles per byte, flat from 512 bytes to 32 KB.
Test the two pointers against each other before going wide, as
arch_strcpy.S already does. Pointers that agree still reach the word
loop, since walking src to a boundary walks dst to one as well; pointers
that disagree take the byte path, where no single boundary serves both.
After the change the misaligned case runs at 490 MB/s and the aligned
rates are unchanged.
The measurements come from the benchmark in apache/nuttx-apps#3706.
Assisted-by: Claude:claude-opus-5
Signed-off-by: Justin Hammond <justin@dynam.ac>
libc_data_t is 8 bytes wide, so a buffer which is 4-byte but not
8-byte aligned falls back to the byte at a time loop. Add a 32-bit
middle path so such buffers still handle four bytes per iteration.
* Add DETECTNULL32/DETECTCHAR32, UNALIGNED4/UNALIGNED4_X,
LITTLEBLOCKSIZE4/BIGBLOCKSIZE4 and TOO_SMALL4 to libs/libc/libc.h.
* Take the new path in memccpy, memcmp, memcpy, memset, stpcpy,
stpncpy, strcmp, strcpy, strncmp and strncpy when both pointers are
4-byte aligned but the 8-byte path can't be used.
Assisted-by: Claude:claude-opus-5
Signed-off-by: Xiang Xiao <xiaoxiang@xiaomi.com>
When the 'c' parameter has bit 7 set (e.g. 0x80), the int value gets
sign extended (to 0xffffff80 on the signed char platforms). The word
sized fill pattern was built without truncating to unsigned char
first, so the fast word aligned path wrote the wrong bytes.
Fix both lib_memset.c and lib_bsdmemset.c by casting 'c' to unsigned
char before building the fill pattern, as required by C11 7.24.6.1
which states that memset converts 'c' to unsigned char.
Assisted-by: Claude:claude-opus-5
Signed-off-by: Bowen Wang <wangbowen6@xiaomi.com>
memrchr scans backward, so the original implementation aligned
(x + 1) rather than x:
#define UNALIGNED(x) ((long)(uintptr_t)((x) + 1) & (sizeof(long) - 1))
while the common UNALIGNED_X() macro checks the pointer itself. Pass
src0 + 1 to UNALIGNED_X() to restore the original behavior, otherwise
asrc is off by one byte and the word loop reads the wrong data.
Assisted-by: Claude:claude-opus-5
Signed-off-by: anjiahao <anjiahao@xiaomi.com>
Remove the incorrect address restoration logic in the memrchr fast
path. The UNALIGNED_X loop already ensures the proper alignment, so
the subsequent address recalculation is unnecessary and makes memrchr
return the wrong position.
This fixes the syslog message corruption where memrchr reports the
incorrect newline position.
Assisted-by: Claude:claude-opus-5
Signed-off-by: fangpeina <fangpeina@xiaomi.com>
Most hardware accesses the memory through a 64-bit bus, so handle the
data in 64-bit chunks instead of "long" chunks which are only 32-bit
wide on the 32-bit platforms.
* Add the libc_data_t type (unsigned long long) and move the shared
UNALIGNED/UNALIGNED_X/ALIGNED, LITTLEBLOCKSIZE, TOO_SMALL and
DETECTNULL helpers from the individual C files to libs/libc/libc.h.
* Convert all lib_bsd*.c implementations to the new type and macros,
which also drops the duplicated LONG_MAX conditionals.
Assisted-by: Claude:claude-opus-5
Signed-off-by: anjiahao <anjiahao@xiaomi.com>
wait4() is BSD/Linux-standard (used by toybox's "time" applet) but NuttX
only had waitpid()+getrusage() separately. Add it to libs/libc/unistd/
built on top of those two existing primitives, so it needs no syscall
plumbing of its own and works unmodified across flat/protected/kernel
build separation. Prototype added to include/sys/wait.h.
Signed-off-by: Alan C. Assis <acassis@gmail.com>
Assisted-by: Claude Sonnet 5 <noreply@anthropic.com>
padlen = sizeof(void *) - (addr % sizeof(void *)) never returns 0, even
when addr is already pointer-aligned -- it returns a full alignment unit
instead. Since callers size buflen for zero padding, the subsequent
"buflen < padlen + reqdlen" check then always fails, so getgrgid()/
getgrnam() and their _r variants always return ERANGE.
Found via `id` on sim:toybox, which resolves gid 0 to "root" through
this path.
Signed-off-by: Alan C. Assis <acassis@gmail.com>
Assisted-by: Claude Sonnet 5 <noreply@anthropic.com>