diff --git a/Documentation/guides/fork_vfork_migration.rst b/Documentation/guides/fork_vfork_migration.rst new file mode 100644 index 00000000000..1124fba1fd9 --- /dev/null +++ b/Documentation/guides/fork_vfork_migration.rst @@ -0,0 +1,215 @@ +========================================= +Migrating to separate ``fork``/``vfork`` +========================================= + +What changed +============ + +NuttX used to implement ``fork()`` and ``vfork()`` as the same function. Both +were thin libc wrappers around a single ``up_fork()`` syscall; ``vfork()`` +differed only by a trailing ``waitpid()``. Underneath, the child joined the +parent's address environment -- the same ``addrenv_join()`` that +``pthread_create()`` uses -- and got a private *copy of the stack*. So the +child shared ``.data``, ``.bss`` and the heap with its parent, and ran +concurrently with it. + +That was not ``fork()``. It was ``vfork()``-with-a-private-stack published +under ``fork()``'s name. The history says so plainly: the ``fork()`` this +replaces was NuttX's old ``vfork()``, renamed in 2023 without any change of +behaviour. And the consequence was silent: a program written against POSIX +``fork()`` compiled and ran, and its child's writes quietly landed in the +parent's variables. + +There are now three distinct primitives: + +.. list-table:: + :header-rows: 1 + :widths: 14 30 26 30 + + * - API + - Memory + - Parent + - Availability + * - ``fork()`` + - child gets **its own copy** at the same virtual addresses + - runs concurrently + - ``CONFIG_ARCH_HAVE_FORK`` -- only where an address environment can be + duplicated + * - ``vfork()`` + - child **shares** the parent's memory + - **suspended** until the child ``_exit()``\ s or ``exec()``\ s + - ``CONFIG_ARCH_HAVE_VFORK`` -- no address environment needed + * - ``task_fork()`` + - child shares memory, private stack copy + - runs concurrently + - ``CONFIG_ARCH_HAVE_TASK_FORK`` -- exactly where ``fork()`` existed + before + +``task_fork()`` is the old behaviour under an honest name. Nothing was lost. + +.. note:: + + ``fork()`` is provided only where the architecture implements + ``up_addrenv_fork()`` and therefore selects + ``CONFIG_ARCH_HAVE_ADDRENV_FORK``; it becomes available architecture by + architecture as that hook lands. Check ``CONFIG_ARCH_HAVE_FORK`` in your + own configuration rather than assuming either way. Where it is unset, + ``vfork()`` and ``task_fork()`` are what the configuration offers, and + ``CONFIG_FORK_IS_TASK_FORK`` keeps the spelling ``fork()`` working for code + that cannot be changed. + +This is a breaking change +========================= + +Two things break, and they break loudly rather than quietly: + +**Code calling** ``fork()`` **on a target without a duplicable address +environment no longer builds.** ``fork()`` is not declared in ``unistd.h`` +there, so you get a compile error naming the function. That is the intended +outcome: a build error is strictly better than the silent wrongness it +replaces. + +**Code calling** ``fork()`` **on a target that does have real** ``fork()`` +**changes behaviour** -- from sharing to copying. Code that (perhaps +unknowingly) relied on the sharing will now see the parent and child diverge. + +Which replacement do I want? +============================ + +Answer the question "why did I call ``fork()``?". + +*I want the child to run a different program.* + Use :c:func:`posix_spawn` or ``task_spawn()``. This is the single most + common reason to call ``fork()``, NuttX has always provided a better answer + for it, and that answer does not have the pid discontinuity that + ``fork()``\ +\ ``exec()`` has. If you must keep the two-step idiom, use + ``vfork()`` + ``exec*()``: that is exactly what ``vfork()`` is for, and + unlike ``fork()`` it needs no duplicable address environment. + +*I want a second flow of control that shares my memory.* + Use ``pthread_create()``. That is the same memory relationship, spelled + clearly, with a normal entry point instead of a function that returns twice. + If you specifically need the returns-twice shape -- for example you are + porting code and do not want to restructure it -- use ``task_fork()``. It + is a rename, not a rewrite: + + .. code-block:: c + + #include + + pid = task_fork(); /* was: pid = fork(); */ + +*I want a genuinely independent copy of this process.* + Keep calling ``fork()``, and make sure your configuration selects + ``CONFIG_ARCH_HAVE_FORK``. Be aware there is no copy-on-write: the copy is + eager, so forking a large process needs as much free memory as the process + occupies and fails with ``ENOMEM`` otherwise. + +*I cannot change the code right now.* + Set ``CONFIG_FORK_IS_TASK_FORK=y``. This aliases ``fork()`` back to + ``task_fork()``, restoring the previous behaviour **exactly** -- same + sharing, same concurrency, no new suspension. It is available on precisely + the configurations that had ``fork()`` before, and it depends on + ``!ARCH_HAVE_FORK``: on a target that can provide real ``fork()``, aliasing + it back to sharing would reintroduce the very ambiguity this change removes, + so sharing-dependent callers there must be edited. + +Configuration symbols +===================== + +``CONFIG_ARCH_HAVE_TASK_FORK`` + Hidden. The architecture can clone the calling task with a copied stack. + Inherits exactly the ``select`` lines that ``ARCH_HAVE_FORK`` used to have, + so no configuration that had ``fork()`` loses the machinery. + +``CONFIG_ARCH_HAVE_VFORK`` + Hidden. The architecture can implement POSIX ``vfork()``. + +``CONFIG_ARCH_HAVE_ADDRENV_FORK`` + Hidden. The architecture implements ``up_addrenv_fork()``, which duplicates + an address environment into freshly allocated pages mapped at the same + virtual addresses. + +``CONFIG_ARCH_HAVE_FORK`` + Hidden, derived: ``ARCH_ADDRENV && ARCH_HAVE_ADDRENV_FORK``. It no longer + means "``fork()`` exists"; it means "this configuration can provide POSIX + ``fork()`` semantics". + +``CONFIG_FORK_IS_TASK_FORK`` + Visible, default ``n``. The legacy alias described above. + +Notes for architecture maintainers +================================== + +The register/stack snapshot machinery is common to all three primitives. Each +architecture exposes three entry points -- ``up_task_fork()``, ``up_vfork()`` +and ``up_fork()`` -- which share one snapshot sequence and differ only in a +``FORK_TYPE_*`` selector (see ``include/nuttx/fork.h``) handed to +``nxtask_setup_fork()``. That is where the memory semantics are decided: +``addrenv_join()`` for ``task_fork()`` and ``vfork()``, ``addrenv_fork()`` for +``fork()``. + +Adding real ``fork()`` to an architecture +----------------------------------------- + +Two things are needed, and the second is the one that is easy to miss. + +**Implement** ``up_addrenv_fork()``. It duplicates an address environment: +allocate fresh pages, copy the parent's contents into them, and map them at the +*same* virtual addresses. ``up_addrenv_clone()`` is not this -- it copies only +the representation and leaves both processes pointing at one set of page +tables. Then give ``ARCH_HAVE_ADDRENV_FORK`` a ``default y if `` line in +``arch/Kconfig``, and ``ARCH_HAVE_FORK`` follows. + +**Build the child from the caller's saved system call frame.** In a kernel +build ``fork()`` is reached through a system call, so the return address and +stack pointer the architecture's fork entry point can observe for itself belong +to the *kernel*, not to the caller; a child built from those resumes at a +kernel address on a kernel stack. The architecture must record the caller's +exception frame when it traps -- ``xcp.sregs`` is the field that exists for +this -- and build the child from that instead, while a kernel thread that calls +the entry point directly still takes the ordinary path. + +Nothing else is required: the ``up_fork()`` entry point and the libc wrapper +are already there and become live automatically. + +Note that ``ARCH_HAVE_ADDRENV_FORK`` is about a *per-process* address +environment. A protected build has one address space carved up once at boot, +whether the boundaries are drawn by an MPU or by a fixed set of MMU mappings; +its ``up_addrenv_*()`` are stubs, and there is no mapping to duplicate at the +same virtual addresses. ``CONFIG_ARCH_ADDRENV`` being set is therefore not by +itself evidence that ``fork()`` can be provided. ``vfork()`` and +``task_fork()``, which share the parent's memory, work there as everywhere +else. + +Known gaps +========== + +``fork()`` **is gained one architecture at a time.** The generic machinery is +complete -- ``addrenv_fork()``, the ``up_addrenv_fork()`` hook, the syscall, the +libc wrapper and the ``ostest`` case -- so an architecture provides ``fork()`` +by implementing ``up_addrenv_fork()`` and selecting +``CONFIG_ARCH_HAVE_ADDRENV_FORK``, with no further generic work. + +**A windowed ABI needs its stack rebased, not just copied.** On Xtensa, +giving a child a relocated copy of the parent's stack takes more than the copy: +the register-window save areas embedded in the stack hold absolute stack +pointers, so each one has to be rebased along with the copy, or the child +reloads a pointer into the *parent's* stack on its very first window underflow. +That rebasing is architecture-specific and belongs with the Xtensa entry points +rather than here. + +Note also on ``waitpid()`` after ``vfork()`` +============================================ + +The ``vfork()`` parent is now resumed when the child's TCB is torn down, so by +the time it runs the child is completely gone. Where the child called +``exec()`` this makes no difference -- ``exec_swap()`` has already given the +loaded program the child's pid, and that program is still running, so +``waitpid()`` behaves normally. Where the child called ``_exit()``, +``waitpid()`` can only return its status if ``CONFIG_SCHED_CHILD_STATUS`` is +enabled; otherwise it returns ``ECHILD``, because NuttX does not retain the +status of a task that no longer exists. That is a pre-existing property of +that configuration, not a change: the previous implementation blocked in a +libc ``waitpid(WNOWAIT)`` and an application's own ``waitpid()`` afterwards hit +the same wall. diff --git a/Documentation/guides/index.rst b/Documentation/guides/index.rst index acbb2a118b5..9b6114409aa 100644 --- a/Documentation/guides/index.rst +++ b/Documentation/guides/index.rst @@ -37,6 +37,7 @@ Guides logging_rambuffer.rst ipv6.rst integrate_newlib.rst + fork_vfork_migration.rst protected_build.rst platform_directories.rst port_drivers_to_stm32f7.rst diff --git a/Documentation/implementation/memory_configurations.rst b/Documentation/implementation/memory_configurations.rst index 6ad536df5d6..eded4ff3096 100644 --- a/Documentation/implementation/memory_configurations.rst +++ b/Documentation/implementation/memory_configurations.rst @@ -30,7 +30,7 @@ On-Demand Paging NuttX also supports on-demand paging via ``CONFIG_PAGING``. On-demand paging is a method of virtual memory management and requires -the the CPU architecutre support a MMU. +the the CPU architecture support a MMU. In a system that uses on-demand paging, the OS responds to a page fault by copying data from some storage media into physical memory and setting up @@ -410,7 +410,7 @@ of functions that: 1. Have only one ``.text`` space in RAM, but 2. Separate ``.data`` and ``.bass`` space, and are -3. Separately linked into with the program in each address environmnet. +3. Separately linked into with the program in each address environment. (not implemented). @@ -484,7 +484,7 @@ at least in its current form. That full implementation of ``mmap()`` plus the minor changes to the NuttX ELF loader are all that are required to support fully share-able ``.text`` sections – as well as the memory savings -from not carrying aroung the relocation and symbol information +from not carrying around the relocation and symbol information (Not implemented). @@ -632,10 +632,13 @@ the contemplate in any real detail: and swap the state into physical memory as needed?(not implemented). * ``mmap()``. True shared memory and true file mapping could be supported. I am repeating myself (not implemented). -* ``fork()``. The ``fork()`` interface could be supported. NuttX currently - supports the "crippled" version, ``vfork()`` but with these process address - environments, the real ``fork()`` interface could be supported. - (not implemented). +* ``fork()``. The real ``fork()`` interface can be supported on configurations + with a duplicable process address environment: an architecture implements + ``up_addrenv_fork()``, selects ``CONFIG_ARCH_HAVE_ADDRENV_FORK``, and + ``CONFIG_ARCH_HAVE_FORK`` follows. What is not implemented is + copy-on-write: the duplication copies the parent's pages eagerly, which + needs as much free memory as the parent occupies. Demand paging would fix + that. * Dynamic Stack Allocation. Completely eliminate the need for constant tuning of static stack sizes.(not implemented). * Shared Libraries. Am I repeating myself again?(not implemented). @@ -688,8 +691,8 @@ There are two problems here: So how do you create new tasks/processes in such a context. There is only one way possible; by using an interface that takes a file name as an argument (rather than absolute address). -New processes started with ``vfork()`` and ``exec()`` or with -``posix_spawn()`` should not have any of these issues. +New processes started with ``fork()`` or ``vfork()`` and ``exec()``, or with +``posix_spawn()``, should not have any of these issues. ARM Memory Management diff --git a/Documentation/reference/user/01_task_control.rst b/Documentation/reference/user/01_task_control.rst index 163030776c3..67b86677148 100644 --- a/Documentation/reference/user/01_task_control.rst +++ b/Documentation/reference/user/01_task_control.rst @@ -59,9 +59,11 @@ Standard interfaces - :c:func:`exit` - :c:func:`getpid` -Standard ``vfork`` and ``exec[v|l]`` interfaces: +Standard ``fork``/``vfork`` and ``exec[v|l]`` interfaces: + - :c:func:`fork` - :c:func:`vfork` + - :c:func:`task_fork` (non-standard) - :c:func:`exec` - :c:func:`execv` - :c:func:`execl` @@ -347,6 +349,47 @@ Functions **POSIX Compatibility:** Compatible with the POSIX interface of the same name. +.. c:function:: pid_t fork(void) + + ``fork()`` creates a new process. The child process is an exact copy of + the calling process: it receives **its own copy** of the parent's memory, + at the same virtual addresses. Writes by the child are invisible to the + parent and writes by the parent are invisible to the child. The child may + modify anything, call any function, return from the function in which + ``fork()`` was called, and run indefinitely; it runs concurrently with the + parent. None of ``vfork()``'s restrictions apply. + + NOTE: ``fork()`` requires an address environment that can be + duplicated, and so it is available only where ``CONFIG_ARCH_HAVE_FORK`` + is selected -- which in turn requires ``CONFIG_ARCH_ADDRENV`` and an + architecture that implements ``up_addrenv_fork()``. **Where it cannot + be provided, ``fork()`` is not provided at all**: the declaration is + absent from ``unistd.h`` and code that calls it fails to build. That + is deliberate. A build error naming the function is strictly better + than a ``fork()`` that silently gives the child the parent's memory. + See :doc:`/guides/fork_vfork_migration` for how to move code that + relied on the previous behaviour. + + There is no copy-on-write, because NuttX has no demand paging to build + it on, so the copy is eager: forking a large process needs as much free + memory as the process occupies and fails with ``ENOMEM`` otherwise. + Spawn-heavy code should prefer :c:func:`posix_spawn` or + :c:func:`vfork`, on NuttX as anywhere. + + Applications that want the historical NuttX ``fork()`` behaviour -- + shared memory, private stack, both running -- should call + :c:func:`task_fork`. ``CONFIG_FORK_IS_TASK_FORK`` aliases ``fork()`` + back to it on configurations that cannot provide real ``fork()``, for + legacy code that cannot be changed. + + :return: Upon successful completion, ``fork()`` returns 0 to the child + process and returns the process ID of the child process to the parent + process. Otherwise, -1 is returned to the parent, no child process is + created, and ``errno`` is set to indicate the error. + + **POSIX Compatibility:** Compatible with the POSIX interface of the same + name. + .. c:function:: pid_t vfork(void) The ``vfork()`` function has the same effect as @@ -357,12 +400,18 @@ Functions function before successfully calling ``_exit()`` or one of the ``exec`` family of functions. - NOTE: ``vfork()`` is not an independent NuttX feature, but is - implemented in architecture-specific logic (using only helper - functions from the NuttX core logic). As a result, ``vfork()`` may - not be available on all architectures. The current implementation in - NuttX arm64 only guarantees that ``vfork()`` works when - CONFIG_BUILD_FLAT=y. + The child **shares** the parent's memory -- nothing is copied, which is the + entire point of ``vfork()`` and the reason a caller chooses it -- and the + parent is **suspended** until the child calls ``_exit()`` or one of the + ``exec`` family. That suspension is what makes the sharing safe, and it is + also the price: the restrictions above exist because the child is running + in the parent's address space on borrowed time. + + NOTE: ``vfork()`` is implementable with or without an MMU and is + available wherever ``CONFIG_ARCH_HAVE_VFORK`` is selected. The + suspension lives in the kernel primitive rather than in a libc + ``waitpid()``, so the parent is resumed at ``exec()`` as POSIX requires, + and ``vfork()`` does not depend on ``CONFIG_SCHED_WAITPID``. :return: Upon successful completion, ``vfork()`` returns 0 to the child process and returns the process ID of the child process to the @@ -372,6 +421,34 @@ Functions **POSIX Compatibility:** Compatible with the BSD/Linux interface of the same name. POSIX marks this interface as Obsolete. +.. c:function:: pid_t task_fork(void) + + ``task_fork()`` is a non-standard NuttX interface that clones the calling + task. The child shares the parent's ``.data``, ``.bss`` and heap -- it + joins the parent's address environment, exactly as a pthread does -- but + runs on a **private copy** of the parent's stack, and it runs concurrently + with the parent. Like ``fork()`` it returns twice. + + This is neither ``fork()`` nor ``vfork()``. It is a task cloned at the + call site with the memory relationship of a thread; the nearest precedent + is Plan 9's ``rfork(RFPROC|RFMEM)``. It is the behaviour NuttX published + under the name ``fork()`` before these three interfaces were separated, and + it is preserved here under an honest name so that nothing is lost. + + NOTE: available where ``CONFIG_ARCH_HAVE_TASK_FORK`` is selected, which + is exactly the set of configurations that had ``fork()`` before the + split. New code should prefer :c:func:`pthread_create`, which is the + same memory relationship spelled clearly, or :c:func:`posix_spawn`; + ``task_fork()`` exists to give the historical behaviour a truthful name, + not to recommend it. + + :return: Upon successful completion, ``task_fork()`` returns 0 to the child + and returns the process ID of the child to the parent. Otherwise, -1 is + returned to the parent, no child is created, and ``errno`` is set to + indicate the error. + + **POSIX Compatibility:** Non-standard. + .. c:function:: int exec(FAR const char *filename, FAR char * const *argv, FAR const struct symtab_s *exports, int nexports) This non-standard, NuttX function is similar to @@ -449,7 +526,15 @@ Functions thread, then (2) call ``execv()`` or ``execl()`` to replace the new thread with a program from the file system. Since the new thread will be terminated by the ``execv()`` or ``execl()`` call, it really served no - purpose other than to support POSIX compatibility. + purpose other than to support POSIX compatibility. :c:func:`posix_spawn` + does the same job in one step and should be preferred. + + Note also that ``exec()`` does not overlay the calling process: it starts + the new program as a separate task. ``exec_swap()`` then exchanges the two + pids, so that from the parent's point of view the pid ``vfork()`` returned + does name the running program -- but the ``vfork()`` stub itself exits, and + it is that exit which releases the suspended ``vfork()`` parent. The parent + therefore resumes at ``exec()``, as POSIX requires. The non-standard binfmt function ``exec()`` needs to have (1) a symbol table that provides the list of symbols exported by the base code, and diff --git a/Documentation/standards/posix.rst b/Documentation/standards/posix.rst index cdca6eb748e..47622ec7921 100644 --- a/Documentation/standards/posix.rst +++ b/Documentation/standards/posix.rst @@ -1421,7 +1421,7 @@ Multiple Processes: +--------------------------------+---------+ | :c:func:`exit` | Yes | +--------------------------------+---------+ -| fork() | No | +| :c:func:`fork` | Cond. | +--------------------------------+---------+ | :c:func:`getpgrp` | Yes | +--------------------------------+---------+ @@ -2364,7 +2364,7 @@ XSI Multiple Process: +--------------------------------+---------+ | :c:func:`usleep` | Yes | +--------------------------------+---------+ -| :c:func:`vfork` | Yes | +| :c:func:`vfork` | Cond. | +--------------------------------+---------+ | :c:func:`waitid` | Yes | +--------------------------------+---------+