Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260904-1856-tyrell-canary-accepted-daemon-still-leaking

rdmsm4x-changelog-20260904-1856-tyrell-canary-accepted-daemon-still-leaking

Deployed and accepted the signed Tyrell leak-fix app canary on rdmbair15m5, and proved it does NOT fix INC-20260904-01: the app pipeline never touches the daemon, which still runs the leaking build. Fleet rollout deliberately NOT started. Undo: rollback bundle retained on the canary; remove three scratch worktrees.

[2026-09-04 18:56:00 EDT · rdmsm4x]

Scope

Host rdmsm4x (Aqua desktop session), plus one remote change on rdmbair15m5 (the designated canary). No change to the other four spokes.

What changed

1. ~/dev/apps/Tyrell — one commit, pushed

060300d scripts(deploy): integrate the plan/redeploy/probe/accept edits the rollout depends on. Four deploy scripts had sat uncommitted on canonical main since 02:15–02:22 EDT, authoring session unknown, holding no ticket claim, while the INC-20260904-01 rollout script called the flags they add. Reviewed hunk by hunk, zsh -n clean, committed by explicit path only, snapshot first to ~/.agent-coordination/snapshots/rdmsm4x/Tyrell/20260904-184045. The canonical-main integrity gate (INCIDENT-20260901-04) was overridden with its own documented TYRELL_CANONICAL_INTEGRATION=1, which exists for reviewed integration-only commits. Pushed to fleet and backup (45e976b..060300d). The Mem0/UI work in progress in that tree was left untouched and is still uncommitted.

2. Clean signed artifact, built in isolation

Three detached worktrees under ~/dev/_scratch/tyrell-main-45e976b-sign-20260904/: Tyrell at 45e976b, ecs0lib at 22888aa, ECSCloudKit at 93b9f70, all dirty=0. Built there rather than in canonical main, whose tree is dirty with another agent's in-flight work.

3. Canary on rdmbair15m5 — deployed and accepted

plan_app_deployment.zsh --candidate-bundle rc=0 · redeploy_app_fleet.zsh --execute-canary rc=0 · run_tyrell_app_canary_probe.zsh rc=0 · accept_tyrell_app_canary.zsh rc=0. Receipt ~/dev/_ops/incidents/20260904-fleet-spawn-starvation-reaper/deploy-runs/fdfix-20260904185210.receipt.json (sha256 1c888c23f0bc4151…), status=accepted, accepted_endpoint=43118, gui/status/usage/chat contracts all passed. Rollback bundle retained on the canary at /Applications/.tyrell-rollbacks/20260904-185238-2af2198ad96d/Tyrell.app (pre-install exec sha256 04a3f96c…, cdhash 2af2198a…).

The finding that matters

After an accepted canary, the canary's daemon is unchanged:

before after
/Applications/Tyrell.app exec sha256 04a3f96c… 8597c9fc… (fixed, SubprocessRunner=196)
LaunchAgent daemon sha256 442f1924… 442f1924… (unchanged)
daemon SubprocessRunner / FileDescriptorLimit / autoreleasePool 0 / 0 / 0 0 / 0 / 0
daemon process pid 1011, 13h33m, RSS 436,944 KB same process, never restarted
launchctl limit maxfiles 256 256

The app-deployment pipeline replaces /Applications/Tyrell.app only. The LaunchAgent runs ~/Library/Application Support/Tyrell/releases/b15-universal/tyrelld, which it never touches. So the premise recorded in SESSION-STATE — that sign_and_stage_after_console_login.zsh --fleet is the one command that fixes the five leaking hosts — is wrong. Run as written it would leave all five daemons leaking while every gate reported green.

Why the fleet stage was NOT run

The daemon's own path is scripts/deploy_fleet.zsh: rsync the repo to each host, swift build -c release --product tyrelld remotely, then install_service.zsh. Three blockers, each needing an owner decision rather than an improvisation:

  1. It rsyncs ~/dev/apps/Tyrell from this host, and canonical main is dirty with another agent's uncommitted work. Running it now pushes unreviewed source to five hosts and rebuilds there. The rsync carries --delete scoped to that directory on the target.
  2. It builds on each target instead of shipping the signed universal binary already verified here, so the two Intel hosts depend on their own toolchain and arch parity is unproven.
  3. It is the developer-build-path shape that ISSUE-20260831-17 exists about.

Recommended: add a daemon-staging path that ships the signed universal tyrelld (dbc440bd54b85a54dd6c4150b23b3e508bff7d8d8080687d81b9ba7ac2cbbcf0, same build) into releases/<tag>/ and repoints the LaunchAgent, matching how the app is already deployed. Gate either route on nm -a "$(plutil -extract ProgramArguments.0 raw ~/Library/LaunchAgents/com.eastcoastscience.tyrelld.plist)" | grep -ci SubprocessRunner being non-zero per host, and on maxfiles rising off 256.

Also done this session

Verification evidence

~/dev/_scratch/tyrell-main-45e976b-sign-20260904/apps/Tyrell/.build/evidence-main-45e976b-20260904-1830/ — REPORT.md, per-file manifest (sha256 3618cde6…), reproducible ustar archive (ba01668f…), redacted build log (a3d1d9f0…, 0 errors, 76 Swift 6 Sendable warnings). Run pinned at deploy-runs/fdfix-20260904185210.env.

Undo

Owner actions outstanding

  1. Decide the daemon rollout route (a) clean the tree and run deploy_fleet.zsh, or (b) ship the signed universal tyrelld into releases/. Five hosts stay leaking until then.
  2. rdmpw3265m and rdmpw3275m still need their own console login before their LaunchAgents can start a daemon at all.
  3. The Mem0/UI work uncommitted in canonical main belongs to another session and is still in the vulnerable verified-and-uncommitted state.