Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260908-1312-tyrelld-leak-remeasured-plateaus-not-runaway

tyrelld leak re-measured across the fleet — it plateaus; nothing is in the tens of GB

Read-only measurement of the tyrelld memory leak on all reachable fleet hosts, plus one documentation commit. Nothing was signed, installed, deployed or restarted; no keychain was touched; ISSUES.md #54 was not run. Requested by the Tyrell cloud oversight thread (session_01CtKKzsMhdkZN1PiY3M4ryc), which has no hands on a Mac and had had no measurement for ~3 days.

Headline

At the 15.0 MB/min hub rate recorded 2026-09-05, an unrestarted tyrelld would now hold 63.3 GB. The largest RSS anywhere on the fleet is 2.35 GB. On three of four probed hosts the measured 30-minute slope is at or below zero. The leak plateaus; it does not run away.

rdmbair13m5 ran 3.6 days unrestarted (pid 48092, runs = 1, never exited) and went 817 MB → 811.8 MB → 576.8 MB across 67.9 hours — net −240 MB. Fleet-wide, RSS is anti-correlated with daemon uptime, which a time-linear leak cannot produce.

Acceptance-metric probe results

scripts/probe_tyrelld_rss_slope.zsh v1.0.0, unmodified, run with defaults. Copied verbatim to /tmp on spokes whose checkouts predate the script; no spoke repo was modified.

rdmsm4x      slope_MB_per_min=-0.510 rss_end_MB=2408 fd_delta=-2 sessions=0 restarts=1 → FAIL
rdmbair15m5  slope_MB_per_min=-0.041 rss_end_MB=461  fd_delta=0  sessions=0 restarts=0 → FAIL
rdmpw3265m   slope_MB_per_min=0.186  rss_end_MB=729  fd_delta=0  sessions=0 restarts=0 → FAIL
rdmpw3275m   slope_MB_per_min=0.934  rss_end_MB=2284 fd_delta=0  sessions=0 restarts=0 → FAIL
rdmbair13m5  slope_MB_per_min=0.000  rss_end_MB=380  fd_delta=0  sessions=0 restarts=0 → FAIL  (re-probed post-reboot, 13:40)
jdmbair13m5  not obtained — unreachable

Every host fails the 300 MB absolute bound; only rdmpw3275m fails the 0.20 MB/min slope rule. What fails is the plateau height, not the rate.

Findings recorded

  1. The probe's restart check is wrong and will fail every host after the rollout. It greps for last exit code = (never exited), which tests whether the job has ever exited, not whether it exited during the window. Installing the signed fix restarts the daemon, so every host will report restarts=1 → FAIL permanently. Verified false-positive on rdmsm4x: pid 3070 ran continuously 11:04:53 → 12:56:39 across the whole window. Script not modified.
  2. tyrelld is not in /Applications/Tyrell.app (which holds only Tyrell and tyrellbar), and it is a different binary on every host — five distinct sha256s, three architecture profiles (two arm64-only, one x86_64-only, two universal), two install layouts. The bundle hash does not attest the daemon.
  3. The deployed bundle is unchanged on all five reachable hosts: CFBundleVersion 7, MacOS/Tyrell sha256 75f4190c1156…, designated requirement 7467a4ae84b1… (CLAUDE.md §6 invariant intact), lipo -archs = x86_64 arm64, TeamIdentifier ZU2882L4HT.
  4. The GitHub mirror was 6 commits AHEAD of canonical, inverting the CLAUDE.md §3 lag. ISSUES.md #54 existed only on the mirror, which is why it was not findable from canonical.
  5. Five of six hosts rebooted today (08:16 – 12:55). Not tyrelld dying — every host reports runs = 1 / never exited since its own boot. Flagged as an adjacent fleet-stability question, not investigated.
  6. jdmbair13m5 not reachable: up on ICMP (2/2) and Tailscale (active; direct), but ssh accept-then-resets (kex_exchange_identification: read: Connection reset by peer). Its fleet telemetry is 11.6 h stale and its last census claimed 151 claude processes vs 0–23 elsewhere — unconfirmed, no shell obtainable.

Files touched

File Change
~/dev/apps/Tyrell/SESSION-STATE.md new top checkpoint prepended (+340 lines)

Nothing else was written. Spoke hosts received only a verbatim copy of the probe script at /tmp/probe_tyrelld_rss_slope.zsh (temp, cleared on reboot); no spoke repo was modified.

Git operations

Ref Before After
local ~/dev/apps/Tyrell b33ca8c 1d2c336
backup/main (GitHub mirror) 3a030a6 1d2c336
fleet/main (git.ecs0.net, canonical) b33ca8c 1d2c336
  1. git merge --ff-only backup/main — fast-forward b33ca8c → 3a030a6, picking up the cloud thread's 6 commits. Done before writing anything, so the checkpoint could not revert the two SESSION-STATE commits already on the mirror.
  2. git commit of SESSION-STATE.md only (path-scoped), via TYRELL_CANONICAL_INTEGRATION=1 — the canonical checkout's integrity gate (INCIDENT-20260901-04) blocks direct commits on main. Documentation-only, one file, invariants verified.
  3. git push backup main and git push fleet main. Canonical was 6 commits stale; the push repaired that. All three refs verified in sync at 1d2c336 afterwards.

Verification evidence

How to undo

The only durable change is one commit. git revert 1d2c336 and push, or reset the branch to 3a030a6 (mirror) / b33ca8c (canonical, though that would re-strand the cloud thread's six commits and ISSUES.md #54 — do not).

Outstanding owner actions

No secrets, credentials, tokens or signed URLs appear in this record.

Addendum 2026-09-08 13:40 EDT

rdmbair13m5 re-probed after its 12:55:36 reboot: slope_MB_per_min=0.000 rss_end_MB=380 fd_delta=0 restarts=0 → FAIL (300 MB bound only). Five of six hosts now have a slope: −0.510, −0.041, 0.000, +0.186, +0.934 — only rdmpw3275m exceeds 0.20 MB/min.

This corrects a prediction I made in the checkpoint's §2: a young, mid-rescan daemon was expected to show climbing RSS, and on this host it did not — flat at ~380 MB within ten minutes. rdmbair13m5 holds the smallest corpus on the fleet (235,660 files / 28.8 GB), so its fill completes in minutes. The rescan explanation survives only for the large-corpus hosts and remains unproven for them.

The useful new number is the before/after on one unchanged host: 380 MB fresh vs ~810 MB after 3.6 days ≈ +0.083 MB/min averaged — real accumulation, bounded, under threshold. Shape: fast fill to a few hundred MB, then slow bounded growth over days, then release under pressure.

Two further commits followed the original push:

Final state: local, backup/main and fleet/main all at f81bb23.