The tyrelld plateau reading was wrong — four same-pid daemons grew across one afternoon
Read-only re-measurement of the tyrelld leak on five
fleet hosts, executing the 2026-09-08 13:10 checkpoint's own next
action. The hypothesis that measurement was written to test has
been falsified, and it was my hypothesis.
- Session: 2026-09-08 20:42 – 21:16 EDT,
claude@rdmsm4x - Scope: rdmsm4x, rdmbair15m5, rdmbair13m5, rdmpw3265m, rdmpw3275m; jdmbair13m5 blocked
- Mode: read-only. Nothing restarted,
installed, deployed or signed; no keychain;
ISSUES.md#54 not run; nothing rebooted and no host cleaned up first.
The result
Elapsed since the first set's final snapshot is 7 h 35 m (13:07:44 → 20:42:47), not ~12 h — the low end of the specified window, and the denominator of every rate below.
The test: if the three hosts that read high at ~1 h old have FALLEN toward 400–850 MB, the plateau reading is confirmed; if they KEPT CLIMBING, it is wrong. None fell. Two climbed hard.
| Host | pid then → now | same? | RSS 13:07 | RSS 20:42 | Δ | rate |
|---|---|---|---|---|---|---|
| rdmsm4x | 3070 → 3070 | YES | 2,408.9 MB | 2,436.5 MB | +27.6 | +0.061 MB/min |
| rdmpw3275m | 1033 → 1033 | YES | 2,291.0 MB | 3,134.1 MB | +843.1 | +1.853 MB/min |
| rdmpw3265m | 872 → 872 | YES | 733.0 MB | 1,187.6 MB | +454.6 | +0.999 MB/min |
| rdmbair13m5 | 1025 → 1025 | YES | 387.3 MB | 622.8 MB | +235.5 | +0.518 MB/min |
| rdmbair15m5 | 1001 → 1012 | NO | 457.4 MB | 721.9 MB | — | excluded |
| jdmbair13m5 | — | — | — | — | — | BLOCKED |
rdmpw3275m at 3,134 MB is the largest tyrelld
RSS this fleet has recorded — larger than the 2,031 MB that
started the alarm on 09-05, and larger than the 2.35 GB reported as the
fleet maximum that afternoon. That supersedes the earlier headline.
What I got wrong
Two of the three arguments for a plateau do not survive:
- "rdmbair13m5 is flat — slope 0.000 over 30 min." Same pid 1025 has grown +235 MB at +0.518 MB/min, over twice the acceptance threshold.
- "RSS is anti-correlated with daemon uptime." An artefact of measuring three hosts ~1 h after a fleet-wide reboot. With all hosts 4–10 h old, it is not.
What survives: the 63.3 GB extrapolation is still wrong. Three days at these rates gives rdmpw3275m ≈ 10.9 GB worst case. Bounded-ish is not bounded, and the gap between 63 GB and 2.35 GB was evidence of a smaller slope, not of a ceiling.
The urgency argument softened at 13:10 should be restored.
New issue #69 — the acceptance gate is unsound
The 30-minute probes were also run (post-#56 build, unmodified, defaults):
rdmsm4x slope=0.283 rss_end=2438 fd_delta=0 restarts=0 → FAIL
rdmbair15m5 slope=0.038 rss_end=714 fd_delta=0 restarts=0 → FAIL
rdmbair13m5 slope=3.602 rss_end=714 fd_delta=-1 restarts=0 → FAIL
rdmpw3265m slope=-0.020 rss_end=1184 fd_delta=0 restarts=0 → FAIL
rdmpw3275m slope=0.153 rss_end=3130 fd_delta=0 restarts=0 → FAIL
They do not agree with the 7.6-hour rates — by 0.08× to 6.95×, and on rdmpw3265m they disagree in sign: a daemon that added 455 MB across the afternoon measured as shrinking in its window.
Rule 2 (slope ≤ 0.20) would have PASSED the
fleet's two fastest-growing daemons today (rdmpw3265m at
−0.020, rdmpw3275m at +0.153). Both fail only rule 3, the absolute 300
MB bound — which stops discriminating as soon as a build is any good.
Filed as #69, with a proposed replacement: a
long-interval same-pid delta rather than a longer continuous sample.
#56 fix confirmed in the field
restarts=0 on all five, including the
hub, which still carries runs = 2,
last terminating signal = Terminated: 15, and zero
last exit code = (never exited) matches — exactly the
condition that produced the false positive on 09-08. The old grep would
have set RESTARTED=1; the pid comparison returned
0 correctly.
Blocked, and not worked around
- jdmbair13m5 — one ssh attempt, timed out on the tailnet IP as expected, not chased and not substituted. Unmeasured for a third consecutive pass.
- rdmbair15m5 — pid changed 1001 → 1012;
last rebootshows a third reboot today (08:16, 13:55, 17:11). Its 30-minute slope is reported (a slope needs one pid); its 7.6-hour delta is not. sessions=0on all five again — still false, still #59. Not evidence the hosts were idle.
Files touched
| File | Change |
|---|---|
~/dev/apps/Tyrell/SESSION-STATE.md |
new top checkpoint (+157 lines) |
~/dev/apps/Tyrell/ISSUES.md |
new #69 |
No source, script or config modified. Spokes received a verbatim
probe copy at /tmp/probe2.zsh only; no spoke repo
touched.
Git
| Ref | Before | After |
|---|---|---|
| local | 2072506 |
3e65a84 |
backup/main (mirror) |
aaa42b7 |
3e65a84 |
fleet/main (canonical) |
2072506 |
3e65a84 |
Fast-forwarded 2072506 → aaa42b7 (11
commits, mostly the new CI workflow) before prepending,
per CLAUDE.md §3. Commit via TYRELL_CANONICAL_INTEGRATION=1
(documentation only). Invariants: SESSION-STATE 2784 → 2941 lines,
headings 89 → 90, zero lost; ISSUES distinct numbers 53
→ 54, zero lost, exactly #69 added.
Outstanding owner actions
- #54 (signing) remains Rich's call and was not run.
- CI's
swift build && swift testjob still needs theECS_LIBS_TOKENrepository secret, which only Rich can mint.scripts/setup_ci_ecs_libs_token.shinstalls it in one command. No token was created, harvested, reused, or that script run by this session. - PR #15 (the #65 RepoBundler fix) is open and uncompiled — untouched here.
- jdmbair13m5 needs console access; three passes, no measurement.
- #69 should be settled before any rollout is judged by the metric — the gate would certify a growing daemon.
No secrets, credentials, tokens or signed URLs appear in this record.