tyrelld leak re-measured across the fleet — it plateaus; nothing is in the tens of GB
Read-only measurement of the tyrelld memory leak on all
reachable fleet hosts, plus one documentation commit. Nothing
was signed, installed, deployed or restarted; no keychain was touched;
ISSUES.md #54 was not run. Requested by the Tyrell
cloud oversight thread (session_01CtKKzsMhdkZN1PiY3M4ryc),
which has no hands on a Mac and had had no measurement for ~3 days.
- Session: 2026-09-08 12:25:33 – 13:15 EDT,
claude@rdmsm4x - Scope: rdmsm4x, rdmbair15m5, rdmbair13m5, rdmpw3265m, rdmpw3275m measured; jdmbair13m5 not reached
- Mode: read-only measurement + one docs commit
Headline
At the 15.0 MB/min hub rate recorded 2026-09-05, an unrestarted
tyrelld would now hold 63.3 GB. The
largest RSS anywhere on the fleet is 2.35 GB. On three
of four probed hosts the measured 30-minute slope is at or below zero.
The leak plateaus; it does not run away.
rdmbair13m5 ran 3.6 days unrestarted
(pid 48092, runs = 1, never exited) and went
817 MB → 811.8 MB → 576.8 MB across 67.9 hours — net −240
MB. Fleet-wide, RSS is anti-correlated with daemon
uptime, which a time-linear leak cannot produce.
Acceptance-metric probe results
scripts/probe_tyrelld_rss_slope.zsh v1.0.0,
unmodified, run with defaults. Copied verbatim to
/tmp on spokes whose checkouts predate the script; no spoke
repo was modified.
rdmsm4x slope_MB_per_min=-0.510 rss_end_MB=2408 fd_delta=-2 sessions=0 restarts=1 → FAIL
rdmbair15m5 slope_MB_per_min=-0.041 rss_end_MB=461 fd_delta=0 sessions=0 restarts=0 → FAIL
rdmpw3265m slope_MB_per_min=0.186 rss_end_MB=729 fd_delta=0 sessions=0 restarts=0 → FAIL
rdmpw3275m slope_MB_per_min=0.934 rss_end_MB=2284 fd_delta=0 sessions=0 restarts=0 → FAIL
rdmbair13m5 slope_MB_per_min=0.000 rss_end_MB=380 fd_delta=0 sessions=0 restarts=0 → FAIL (re-probed post-reboot, 13:40)
jdmbair13m5 not obtained — unreachable
Every host fails the 300 MB absolute bound; only rdmpw3275m fails the 0.20 MB/min slope rule. What fails is the plateau height, not the rate.
Findings recorded
- The probe's restart check is wrong and will fail every host
after the rollout. It greps for
last exit code = (never exited), which tests whether the job has ever exited, not whether it exited during the window. Installing the signed fix restarts the daemon, so every host will reportrestarts=1 → FAILpermanently. Verified false-positive on rdmsm4x: pid 3070 ran continuously 11:04:53 → 12:56:39 across the whole window. Script not modified. tyrelldis not in/Applications/Tyrell.app(which holds onlyTyrellandtyrellbar), and it is a different binary on every host — five distinct sha256s, three architecture profiles (two arm64-only, one x86_64-only, two universal), two install layouts. The bundle hash does not attest the daemon.- The deployed bundle is unchanged on all five reachable
hosts:
CFBundleVersion7,MacOS/Tyrellsha25675f4190c1156…, designated requirement7467a4ae84b1…(CLAUDE.md §6 invariant intact),lipo -archs=x86_64 arm64, TeamIdentifierZU2882L4HT. - The GitHub mirror was 6 commits AHEAD of canonical,
inverting the CLAUDE.md §3 lag.
ISSUES.md#54 existed only on the mirror, which is why it was not findable from canonical. - Five of six hosts rebooted today (08:16 – 12:55).
Not
tyrellddying — every host reportsruns = 1/never exitedsince its own boot. Flagged as an adjacent fleet-stability question, not investigated. - jdmbair13m5 not reachable: up on ICMP (2/2) and
Tailscale (
active; direct), but ssh accept-then-resets (kex_exchange_identification: read: Connection reset by peer). Its fleet telemetry is 11.6 h stale and its last census claimed 151claudeprocesses vs 0–23 elsewhere — unconfirmed, no shell obtainable.
Files touched
| File | Change |
|---|---|
~/dev/apps/Tyrell/SESSION-STATE.md |
new top checkpoint prepended (+340 lines) |
Nothing else was written. Spoke hosts received only a verbatim copy
of the probe script at /tmp/probe_tyrelld_rss_slope.zsh
(temp, cleared on reboot); no spoke repo was modified.
Git operations
| Ref | Before | After |
|---|---|---|
local ~/dev/apps/Tyrell |
b33ca8c |
1d2c336 |
backup/main (GitHub mirror) |
3a030a6 |
1d2c336 |
fleet/main (git.ecs0.net, canonical) |
b33ca8c |
1d2c336 |
git merge --ff-only backup/main— fast-forwardb33ca8c→3a030a6, picking up the cloud thread's 6 commits. Done before writing anything, so the checkpoint could not revert the two SESSION-STATE commits already on the mirror.git commitofSESSION-STATE.mdonly (path-scoped), viaTYRELL_CANONICAL_INTEGRATION=1— the canonical checkout's integrity gate (INCIDENT-20260901-04) blocks direct commits onmain. Documentation-only, one file, invariants verified.git push backup mainandgit push fleet main. Canonical was 6 commits stale; the push repaired that. All three refs verified in sync at1d2c336afterwards.
Verification evidence
- Structural invariants on
SESSION-STATE.md: 2091 → 2431 lines,##headings 85 → 86, and zero prior headings lost, checked withcommagainstgit show HEAD:SESSION-STATE.md. - Designated requirement hash recomputed the way
make_app_bundle.zsh:215-217computes it (sha256 of thedesignated =>expression string) on all five hosts →7467a4ae84b1…on each. - Restart false-positive proven against
ps -o lstartprocess identity, not against the launchd text the probe reads. - Post-push:
git rev-parseon local,backup/mainandfleet/mainall return1d2c336, and the checkpoint is the first##heading in the pushed file.
How to undo
The only durable change is one commit.
git revert 1d2c336 and push, or reset the branch to
3a030a6 (mirror) / b33ca8c (canonical, though
that would re-strand the cloud thread's six commits and
ISSUES.md #54 — do not).
Outstanding owner actions
ISSUES.md#54 (signing / Aqua-session hypothesis) remains Rich's call and was not run.- jdmbair13m5 needs console access or a reboot to restore ssh; its leak state is unknown.
- Re-measure the five hosts in 6–12 h with no reboot in between to confirm the plateau reading. Read-only, needs no approval.
- The hub's −0.510 slope was taken at
sessions=0, so the acceptance metric's "≥ 10 agent sessions" condition was not met and it is not yet a trustworthy pass.
No secrets, credentials, tokens or signed URLs appear in this record.
Addendum 2026-09-08 13:40 EDT
rdmbair13m5 re-probed after its 12:55:36 reboot:
slope_MB_per_min=0.000 rss_end_MB=380 fd_delta=0 restarts=0 → FAIL
(300 MB bound only). Five of six hosts now have a slope: −0.510,
−0.041, 0.000, +0.186, +0.934 — only rdmpw3275m exceeds 0.20
MB/min.
This corrects a prediction I made in the checkpoint's §2: a young, mid-rescan daemon was expected to show climbing RSS, and on this host it did not — flat at ~380 MB within ten minutes. rdmbair13m5 holds the smallest corpus on the fleet (235,660 files / 28.8 GB), so its fill completes in minutes. The rescan explanation survives only for the large-corpus hosts and remains unproven for them.
The useful new number is the before/after on one unchanged host: 380 MB fresh vs ~810 MB after 3.6 days ≈ +0.083 MB/min averaged — real accumulation, bounded, under threshold. Shape: fast fill to a few hundred MB, then slow bounded growth over days, then release under pressure.
Two further commits followed the original push:
- The cloud thread filed
ISSUES.md#55 (notyrelldin the bundle; five hosts, five daemon binaries) and #56 (the probe's restart check will fail every host after the rollout) from this session's findings, and corrected its own CLAUDE.md §3 claim. Fast-forwarded in before appending. f81bb23— this addendum.
Final state: local, backup/main and
fleet/main all at
f81bb23.