Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260908-2115-tyrelld-plateau-refuted-remeasurement

The tyrelld plateau reading was wrong — four same-pid daemons grew across one afternoon

Read-only re-measurement of the tyrelld leak on five fleet hosts, executing the 2026-09-08 13:10 checkpoint's own next action. The hypothesis that measurement was written to test has been falsified, and it was my hypothesis.

The result

Elapsed since the first set's final snapshot is 7 h 35 m (13:07:44 → 20:42:47), not ~12 h — the low end of the specified window, and the denominator of every rate below.

The test: if the three hosts that read high at ~1 h old have FALLEN toward 400–850 MB, the plateau reading is confirmed; if they KEPT CLIMBING, it is wrong. None fell. Two climbed hard.

Host pid then → now same? RSS 13:07 RSS 20:42 Δ rate
rdmsm4x 3070 → 3070 YES 2,408.9 MB 2,436.5 MB +27.6 +0.061 MB/min
rdmpw3275m 1033 → 1033 YES 2,291.0 MB 3,134.1 MB +843.1 +1.853 MB/min
rdmpw3265m 872 → 872 YES 733.0 MB 1,187.6 MB +454.6 +0.999 MB/min
rdmbair13m5 1025 → 1025 YES 387.3 MB 622.8 MB +235.5 +0.518 MB/min
rdmbair15m5 1001 → 1012 NO 457.4 MB 721.9 MB — excluded
jdmbair13m5 — — — — — BLOCKED

rdmpw3275m at 3,134 MB is the largest tyrelld RSS this fleet has recorded — larger than the 2,031 MB that started the alarm on 09-05, and larger than the 2.35 GB reported as the fleet maximum that afternoon. That supersedes the earlier headline.

What I got wrong

Two of the three arguments for a plateau do not survive:

What survives: the 63.3 GB extrapolation is still wrong. Three days at these rates gives rdmpw3275m ≈ 10.9 GB worst case. Bounded-ish is not bounded, and the gap between 63 GB and 2.35 GB was evidence of a smaller slope, not of a ceiling.

The urgency argument softened at 13:10 should be restored.

New issue #69 — the acceptance gate is unsound

The 30-minute probes were also run (post-#56 build, unmodified, defaults):

rdmsm4x      slope=0.283  rss_end=2438 fd_delta=0  restarts=0 → FAIL
rdmbair15m5  slope=0.038  rss_end=714  fd_delta=0  restarts=0 → FAIL
rdmbair13m5  slope=3.602  rss_end=714  fd_delta=-1 restarts=0 → FAIL
rdmpw3265m   slope=-0.020 rss_end=1184 fd_delta=0  restarts=0 → FAIL
rdmpw3275m   slope=0.153  rss_end=3130 fd_delta=0  restarts=0 → FAIL

They do not agree with the 7.6-hour rates — by 0.08× to 6.95×, and on rdmpw3265m they disagree in sign: a daemon that added 455 MB across the afternoon measured as shrinking in its window.

Rule 2 (slope ≤ 0.20) would have PASSED the fleet's two fastest-growing daemons today (rdmpw3265m at −0.020, rdmpw3275m at +0.153). Both fail only rule 3, the absolute 300 MB bound — which stops discriminating as soon as a build is any good. Filed as #69, with a proposed replacement: a long-interval same-pid delta rather than a longer continuous sample.

#56 fix confirmed in the field

restarts=0 on all five, including the hub, which still carries runs = 2, last terminating signal = Terminated: 15, and zero last exit code = (never exited) matches — exactly the condition that produced the false positive on 09-08. The old grep would have set RESTARTED=1; the pid comparison returned 0 correctly.

Blocked, and not worked around

Files touched

File Change
~/dev/apps/Tyrell/SESSION-STATE.md new top checkpoint (+157 lines)
~/dev/apps/Tyrell/ISSUES.md new #69

No source, script or config modified. Spokes received a verbatim probe copy at /tmp/probe2.zsh only; no spoke repo touched.

Git

Ref Before After
local 2072506 3e65a84
backup/main (mirror) aaa42b7 3e65a84
fleet/main (canonical) 2072506 3e65a84

Fast-forwarded 2072506 → aaa42b7 (11 commits, mostly the new CI workflow) before prepending, per CLAUDE.md §3. Commit via TYRELL_CANONICAL_INTEGRATION=1 (documentation only). Invariants: SESSION-STATE 2784 → 2941 lines, headings 89 → 90, zero lost; ISSUES distinct numbers 53 → 54, zero lost, exactly #69 added.

Outstanding owner actions

No secrets, credentials, tokens or signed URLs appear in this record.