netutils/ptpd: discard implausible drift-rate samples before averaging

A single drift-rate sample computed between two consecutive sync
updates was clamped against CLOCK_ADJTIME_SLEWLIMIT_PPM - the
hardware's slew-rate safety limit, not a bound on how large a real
crystal-oscillator drift measurement can plausibly be. An abnormally
short or long measurement interval (e.g. right after a clock
source outage/reconnect, or a burst of closely spaced sync packets
following packet loss) could therefore produce a wildly implausible
sample that still passed the check and corrupted the long-term
drift_ppb average.

Add CONFIG_NETUTILS_PTPD_MAX_DRIFT_PPB (default 500000, well above
any real crystal's few-hundred-ppm drift) as a dedicated plausibility
bound, intentionally much tighter than CLOCK_ADJTIME_SLEWLIMIT_PPM.
A sample outside this bound is discarded and the previous averaged
drift_ppb is kept unchanged instead of being corrupted.

Assisted-by: Claude:claude-sonnet-5
Signed-off-by: Daniel P. Carvalho <danieloak@gmail.com>
This commit is contained in:
Daniel P. Carvalho 2026-09-16 15:19:41 -03:00 • committed by Xiang Xiao
parent 6d6ab39939
commit e8fa49dc72
2 changed files with 32 additions and 3 deletions

View file

@ -175,6 +175,26 @@ config NETUTILS_PTPD_DRIFT_AVERAGE_S
gives more stable estimate but reacts slower to crystal oscillator speed
changes (such as caused by temperature changes).
config NETUTILS_PTPD_MAX_DRIFT_PPB
int "PTP client maximum plausible clock drift rate (ppb)"
default 500000
range 1000 20000000
---help---
A single drift-rate sample computed between two consecutive sync
updates is discarded (the previous averaged drift_ppb is kept
unchanged) if its magnitude exceeds this bound. Real crystal
oscillators drift by at most a few hundred ppm (hundreds of
thousands of ppb), so this catches bogus samples caused by an
abnormally short or long measurement interval - e.g. right after
a clock source outage/reconnect, or a burst of closely spaced
sync packets following packet loss - before they corrupt the
long-term drift_ppb average and get applied to the hardware.
This is intentionally much tighter than
CLOCK_ADJTIME_SLEWLIMIT_PPM, which bounds how fast a correction
may be applied rather than how large a real drift measurement
can plausibly be.
config NETUTILS_PTPD_MAX_PATH_DELAY_NS
int "PTP client maximum path delay (ns)"
default 100000

View file

@ -1232,8 +1232,6 @@ static int ptp_update_local_clock(FAR struct ptp_state_s *state,
const int64_t max_adjust_ns =
(int64_t)CONFIG_CLOCK_ADJTIME_SLEWLIMIT_PPM *
CONFIG_CLOCK_ADJTIME_PERIOD_MS;
const int64_t slew_limit_ppb =
(int64_t)CONFIG_CLOCK_ADJTIME_SLEWLIMIT_PPM * 1000;
if (!state->has_last_delta)
{
@ -1269,8 +1267,19 @@ static int ptp_update_local_clock(FAR struct ptp_state_s *state,
interval_ms = 1;
}
if (drift_ppb > slew_limit_ppb || drift_ppb < -slew_limit_ppb)
if (drift_ppb > CONFIG_NETUTILS_PTPD_MAX_DRIFT_PPB ||
drift_ppb < -CONFIG_NETUTILS_PTPD_MAX_DRIFT_PPB)
{
/* Physically implausible for a real crystal oscillator -
* almost always the result of an abnormally short interval
* between samples (e.g. a burst of packets right after a
* clock source outage/reconnect) rather than actual drift.
* Discard it instead of letting it corrupt the long-term
* average; CLOCK_ADJTIME_SLEWLIMIT_PPM is a much looser
* hardware safety bound and would let this through
* unchanged.
*/
ptpwarn("Drift estimate out of range: %lld\n",
(long long)drift_ppb);
drift_ppb = state->drift_ppb;