A stalled server no longer reads as "your clock is ahead" #251

Merged
jared merged 3 commits from clock-skew-lag into lotus 2026-09-28 22:40:14 -04:00
3 Commits
Author SHA1 Message Date
Lotus CI f0865115a4 Merge remote-tracking branch 'origin/lotus' into clock-skew-lag
CI / Build & Quality Checks (pull_request) Successful in 3m7s
CI / Trigger Desktop Build (pull_request) Skipped
CI / Docker image build & smoke test (pull_request) Skipped
CI / Secret scan (gitleaks) (pull_request) Successful in 7s
CI / Playwright smoke (e2e) (pull_request) Successful in 10m59s
2026-09-28 22:11:31 -04:00
Lotus CIandClaude Opus 5.5 3e5fdd0dab test(e2e): clock-ahead warning needs a minute of samples (#158)
CI / Build & Quality Checks (pull_request) Successful in 2m57s
CI / Trigger Desktop Build (pull_request) Skipped
CI / Docker image build & smoke test (pull_request) Skipped
CI / Secret scan (gitleaks) (pull_request) Successful in 10s
CI / Playwright smoke (e2e) (pull_request) Canceled after 0s
"Ahead" is now reported only once it has held for a minute of fresh
samples (a stalled server delivers late and reads as ahead). The test sends
its ticks, checks nothing is shown yet, fast-forwards the page clock past a
minute and sends two more.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PPmy3tPq869XDW4njjVaKA
2026-09-28 22:08:13 -04:00
Lotus CIandClaude Opus 5.5 bb569d69a2 fix: a stalled server no longer reads as "your clock is ahead"
CI / Build & Quality Checks (pull_request) Successful in 3m1s
CI / Trigger Desktop Build (pull_request) Skipped
CI / Docker image build & smoke test (pull_request) Skipped
CI / Secret scan (gitleaks) (pull_request) Successful in 7s
CI / Playwright smoke (e2e) (pull_request) Failing after 11m28s
Incident 2026-09-29: the homeserver's host ran out of memory and stalled for
~2 minutes. The /sync that finally went out carried events whose `age` was
computed ~30 s before it arrived, so every client showed "Your computer's
clock is 30 seconds ahead of the server" while the real problem was the
server (all host clocks were within 0.25 s the whole evening).

The skew estimate was the median of the last 5 samples, and a sample is
local skew + delivery delay, so one late /sync with a handful of events
tripped it.

- Estimate = the LOWEST sample of the last 5 minutes: delay only ever adds,
  so the fastest-delivered event is the truest.
- "Behind" (which a delay can't cause) is reported as soon as there are 3
  samples, like before. "Ahead" must hold across samples received at least
  a minute apart, so a single late burst never trips it.
- Samples are aged on the monotonic clock, and a change of the local clock
  (someone fixing it) resets the measurement, so the warning clears at once.
- Only events stamped by our own homeserver are sampled: a federated event's
  origin_server_ts is the other server's clock.
- Wording: "This device's clock is … Voice calls and encrypted messages can
  fail until it's corrected." / call bar "Device clock … : calls may fail"
  (was "will fail").

Unit tests: the incident (late burst after normal traffic, and a fresh
client whose first samples are all late), mixed slow/fast deliveries,
ahead only after a minute, behind at once, hysteresis, clock fixed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PPmy3tPq869XDW4njjVaKA
2026-09-28 21:40:02 -04:00