nuttx/Documentation/os/scheduling/index.rst
Vinicius May f6ecf80ebb Documentation: brand new layout for NuttX documentation.
The documentation grew one page at a time, so the tree follows the
history of who wrote what and not the shape of NuttX. Scheduling is
spread over three places, a driver page can sit above the subsystem
that owns it, and the front page lists everything at the same level.
That is a lot to face when all you want to know is where the scheduler
lives.

This change files every page under the code it describes. It is a move,
not a rewrite: outside the ten pages named below, every page keeps the
text that is already in master, and no page's text is deleted.

What it does:

* Groups the table of contents into nine chapters.
* Moves the OS subsystems under os/: scheduling, memory, drivers,
  filesystem, networking, IPC, interrupts, libs, time.
* Renames the platform pages to the names the source tree uses, and
  derives their tags from the tree instead of by hand.
* Splits guides/ by subject.
* Adds Documentation/redirects.py, with a rule for every page that left
  its old path, so old URLs keep working. The redirect page also carries
  a link's #anchor across to the new page.

Ten pages have text that is new or rewritten. Nine of them are the
landing page of a chapter, which has to exist for the new structure:

    index                  the front page
    os/index               OS Design
    os/scheduling/index    Scheduling
    os/interrupts/index    Interrupts
    os/ipc/index           IPC
    os/time/index          Time and timers
    about/index            About
    developing/index       Developing NuttX
    ReleaseNotes/index     Release notes

The tenth is os/libs/libbuiltin, the only page here with technical
content: libs/libbuiltin/ had no page at all. Five SVG diagrams come
with these pages, hand-written XML with no editor metadata.

Nothing outside Documentation/ is touched.

How it was checked:

* Sphinx builds with -W: no warnings, and no document left outside a
  toctree.
* A script, offered in the PR, proves the narrow claim this rests on.
  For every page outside the ten named above it erases what a move
  touches -- link target, path, tag line, toctree block, table border --
  from the whole old text and the whole new text, and requires the two
  to be byte for byte identical. It also requires every sentence of a
  deleted page to turn up somewhere, and every page that left its old
  path to have a redirect, from a URL that existed, to where its content
  went. It exits non-zero and names the page if any of that is not true,
  and it tests added pages too, so forgetting to declare one cannot make
  it pass.
* An independent audit checked 133 factual claims on these ten pages
  against the tree, one shell command per claim: 130 confirmed, 1
  refuted and fixed here, 2 not checkable.
* tools/checkpatch.sh is clean over the range.

The diff is large because moving a page changes every link that points
to it. Most of it is pure renames, and board pages that gained one tag
line.

Assisted-by: Claude:claude-opus-5
2026-10-08 01:40:54 +08:00

239 lines
9.7 KiB
ReStructuredText

==========
Scheduling
==========
NuttX decides which thread runs, and when. This page describes that
decision: the states a thread moves through, the policies that order threads
of equal priority, and the interface an application uses to choose between
them.
The code lives under ``sched/``.
Priority comes first
====================
NuttX is a **strict priority** scheduler. The thread that runs is always
the ready-to-run thread with the highest priority. A lower-priority thread
runs only while no higher-priority thread is ready; the moment a
higher-priority one becomes ready, it takes the CPU. NuttX is fully
pre-emptible, so that happens immediately, including from inside an
interrupt handler.
Priority alone does not say what happens when two ready threads share the
same priority. That is what the scheduling policies decide, and it is the
only thing they decide.
Thread states
=============
Every thread is in exactly one state, held in its task control block. The
states are defined by ``enum tstate_e`` in ``include/nuttx/sched.h``, and the
comment beside each one sorts it into a group: ``INVALID`` for a TCB that has
not been initialised, ``READY_TO_RUN``, and ``BLOCKED``.
.. figure:: task_states.svg
:align: center
:width: 100%
:alt: A task is created inactive, becomes ready to run, is scheduled onto a
CPU, blocks waiting for a resource, and finally exits.
The states of ``enum tstate_e``, and what moves a thread between them.
``INACTIVE`` carries the ``BLOCKED`` comment as well, though it is not
waiting on anything: it is a task that has been created and not yet
activated.
Two ready-to-run states are left off the diagram to keep it readable, and
they are not alike. ``ASSIGNED`` is ``READYTORUN`` with a CPU already
picked, and it is the conditional one -- it exists only under
``CONFIG_SMP``. ``PENDING`` is always compiled, and it is narrower than its
name suggests: a thread goes there only if it became ready while another
thread held ``sched_lock()`` **and** would have pre-empted it.
``nxsched_add_readytorun()`` tests both halves --
``nxsched_islocked_tcb(rtcb)`` and a new priority higher than the running
thread's -- so a thread that becomes ready at equal or lower priority under
the lock joins the ordinary ready-to-run list instead.
A thread in any *ready-to-run* state is runnable; only one per CPU is
``RUNNING``. A thread in any *blocked* state is waiting for something
specific, and the state says what: a semaphore, a signal, an event, a
message arriving, or room to put one. ``STOPPED`` waits for ``SIGCONT``
under ``CONFIG_SIG_SIGSTOP_ACTION``, and one more state is off the diagram
too: ``WAIT_PAGEFILL``, under ``CONFIG_LEGACY_PAGING``. Sleeping is not a
state of its own -- ``TSTATE_SLEEPING`` is defined as another name for
``TSTATE_WAIT_SIG``. This is why a stack dump tells you not just that a
thread is stuck but what it is stuck on.
Scheduling policies
===================
The policy applies **between threads of equal priority**. It never lets a
lower-priority thread run ahead of a higher-priority one.
.. figure:: policies.svg
:align: center
:width: 100%
:alt: Under SCHED_FIFO a thread keeps the CPU until it blocks; under
SCHED_RR threads of equal priority take turns; under
SCHED_SPORADIC a thread runs at a high priority while it has
budget and drops to a low one when the budget is spent.
The same two threads under each policy, and what the sporadic parameters
mean.
``SCHED_FIFO`` -- run to completion
-----------------------------------
A thread runs until it blocks, exits, or is pre-empted by something of
higher priority. Two threads of the same priority do not share the CPU: the
first one to start keeps it until it gives it up.
This is the most predictable policy and the cheapest one, and it is what a
new task gets -- but only while round robin is switched off.
``nxthread_setup_scheduler()`` picks the policy for every new TCB with a
compile-time test, not a runtime one: ``TCB_FLAG_SCHED_RR`` and a timeslice
of ``CONFIG_RR_INTERVAL`` when that value is positive,
``TCB_FLAG_SCHED_FIFO`` otherwise. So ``CONFIG_RR_INTERVAL`` is not a
per-thread opt-in; setting it changes the default policy of the whole
system.
``SCHED_RR`` -- take turns
--------------------------
Same as ``SCHED_FIFO``, except that a thread which has been running for
``CONFIG_RR_INTERVAL`` milliseconds is moved to the back of the queue of
threads at its own priority, and the next one runs.
``CONFIG_RR_INTERVAL`` counts milliseconds and defaults to 0, which disables
the policy entirely -- and *entirely* is literal: at 0 the ``SCHED_RR`` case
is not compiled into ``nxsched_set_scheduler()`` at all, so asking for the
policy at run time fails rather than being ignored. Set it to a positive
value and, as above, every task starts out round-robin. Use it when several
threads of the same priority have to make progress together, and none of
them blocks often enough to give the others a chance on its own.
``SCHED_SPORADIC`` -- a budget per period
-----------------------------------------
Enabled by ``CONFIG_SCHED_SPORADIC``. This one is worth understanding
before reaching for it, because it behaves unlike the other two.
A sporadic thread has **two** priorities and a **budget**:
.. list-table::
:header-rows: 1
:widths: 34 66
* - ``struct sched_param`` field
- Meaning
* - ``sched_priority``
- The high priority. The thread runs at this priority while it still
has budget left.
* - ``sched_ss_low_priority``
- The low priority. The thread drops to this once the budget is spent.
* - ``sched_ss_init_budget``
- How much CPU time the thread may spend at the high priority. A
``struct timespec``.
* - ``sched_ss_repl_period``
- The length of one cycle, budget included. Also a ``struct
timespec``.
* - ``sched_ss_max_repl``
- How many replenishments may be pending at once
(``CONFIG_SCHED_SPORADIC_MAXREPL``).
While the thread has budget it competes at ``sched_priority``. When the
budget runs out it is demoted to ``sched_ss_low_priority``, so it keeps
running only if nothing else wants the CPU. Then the cycle repeats.
The period *contains* the budget rather than following it, which is the part
worth reading twice. ``sched_sporadic.c`` asserts
``repl_period >= budget`` and spends the difference at the low priority::
remainder = sporadic->repl_period - sporadic->budget;
Give a thread a 10 ms budget and a 100 ms period and it runs high for up to
10 ms, low for the other 90, and is promoted again 100 ms after the cycle
began -- not 100 ms after the budget ran out. ``sched_ss_max_repl`` bounds a
second mechanism rather than this one: when the thread is pre-empted part-way
through its budget, each fragment is replenished one period after it was
consumed, and the field caps how many such replenishments may be outstanding
at once.
What this buys you is a **bounded** amount of high-priority CPU time for a
thread whose workload you do not fully trust: an event handler that is
usually short but occasionally is not. It gets to respond quickly, and it
cannot starve the rest of the system if it misbehaves. A thread that would
otherwise have to be given a low priority -- and therefore a poor response
time -- can be given a high one safely.
Choosing a policy
=================
.. list-table::
:header-rows: 1
:widths: 22 78
* - Policy
- Use it when
* - ``SCHED_FIFO``
- The normal case. Threads are ordered by priority and each runs until
it blocks.
* - ``SCHED_RR``
- Several threads sit at the same priority and all have to progress,
for example a set of equivalent workers.
* - ``SCHED_SPORADIC``
- A thread needs a fast response but its running time is not bounded,
and starving lower-priority work is not acceptable.
``SCHED_OTHER`` and ``SCHED_NORMAL`` exist for portability and are both 0;
``SCHED_NORMAL`` is defined as an alias of ``SCHED_OTHER``, which the header
describes as mapping to ``SCHED_FIFO`` or ``SCHED_RR``. In the code that
mapping goes to ``SCHED_RR`` and only while ``CONFIG_RR_INTERVAL`` is
positive: ``nxsched_set_scheduler()`` compiles ``case SCHED_OTHER`` next to
``case SCHED_RR`` under that same test.
``include/sched.h`` also defines ``SCHED_BATCH`` and ``SCHED_IDLE``. Neither
name appears anywhere else in the tree, and ``nxsched_set_scheduler()`` does
not accept either, so asking for one fails. They are numbers reserved in a
header, not policies you can select.
Application interface
=====================
The POSIX interface to all of the above -- ``sched_setscheduler()``,
``sched_setparam()``, ``sched_yield()``, ``sched_rr_get_interval()`` and the
rest -- is documented in :doc:`/reference/user/02_task_scheduling`.
Two interfaces are specific to NuttX and worth naming here:
``sched_lock()`` / ``sched_unlock()``
Hold off the scheduler without disabling interrupts. Interrupts still
run; what is suspended is the switch to another thread. A thread that
becomes ready while pre-emption is locked goes to ``PENDING``, provided it
would have pre-empted the thread holding the lock. This is
cheaper and far less disruptive than disabling interrupts, and it is
almost always the right tool when the goal is "do not switch away from
me" rather than "do not interrupt me". See
:doc:`preemption_latency` for what each choice costs.
``sched_setaffinity()`` / ``sched_getaffinity()``
Restrict a thread to a set of CPUs under ``CONFIG_SMP``. See
:doc:`smp`.
In this section
===============
.. toctree::
:maxdepth: 1
nuttx_tasking.rst
tasks_vs_threads.rst
kernel_threads_vs_pthreads.rst
processes_vs_tasks.rst
context_switches.rst
preemption_latency.rst
cancellation_points.rst
smp.rst
wqueue.rst
tls.rst
user_identity.rst