Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260901-1853-tyrell-fleet-daemon-deploy-and-head-regression-fix

Tyrell fleet daemon deploy + HEAD regression fix

Host: rdmsm4x · When: 2026-09-01 18:24–18:53 EDT · Session: 9b2e2f83 · Task: TASK-20260831-30

One-line summary: deployed the Tyrell CLI/daemon plane to all 5 spokes and found + fixed a regression on canonical main that would have turned that deploy into a fleet-wide daemon outage.

Scope

Hosts changed: rdmbair15m5, rdmbair13m5, jdmbair13m5, rdmpw3265m, rdmpw3275m. rdmsm4x: source + tickets only — its running daemon was deliberately left on its pinned release.

What changed and why it matters

tyrelld built from HEAD exited with code 0 immediately, in both server and client roles. Commit 03ba9a5 replaced the blocking HTTPServer with ecs0lib's ECS0HTTPServer, whose start() resolves as soon as its NWListener is .ready and returns the bound port. The call site _ = try await dataPlane.value was carried across unchanged — it compiles, binds, prints the banner, then main() returns. launchd reads a CLEAN exit and respawns, so the daemon flaps and writes nothing to the error log. Nothing was running HEAD only because rdmsm4x runs a pinned release (releases/7662e97, 2026-08-28) predating the migration.

Files touched (canonical, commit 5765dc7)

Committed via linked worktree ~/dev/_worktrees/tyrell-daemon-park-deploy-deps-20260901 (branch claude/tyrell-daemon-park-deploy-deps-20260901), fast-forwarded into main — canonical Tyrell main is integration-only after INCIDENT-20260901-04.

Commands run

Verification evidence

All 6 hosts: tyrell + tyrelld + tyrell-mcp present at the resolved bin path, daemon up, /api/status OK. Serving-binary sha equals freshly built sha on all 5 spokes (arm64 78a585221b45; Intel by design differs per host — rdmpw3265m c0f9d09ad6c8, rdmpw3275m d56fd73976b3). launchd runs=1 on the four arm64 hosts; runs=4 on the two Intel hosts from the port handover, both stable since on a fixed pid with advancing uptime.

Daemon fix verified BOTH directions: it now stays up (runs=1, "never exited"), AND still dies loudly against an already-bound port (POSIXErrorCode(rawValue: 48): Address already in use, non-zero exit) — the park did not convert a genuine error into a silent success.

Backups / how to undo

Tickets

Resolved: TASK-20260831-04, ISSUE-20260831-09, ISSUE-20260901-09, ISSUE-20260901-14. Opened: ISSUE-20260901-16 (deploy grades a host green after a FAILED build by bootstrapping the stale binary — observed live, filed rather than fixed silently). Updated: TASK-20260831-30 (partial), ISSUE-20260901-10 (cross-reference only; a commit message mis-cites it).

OUTSTANDING — needs Rich

Tyrell.app is NOT redeployed. ISSUE-20260831-17's fix is built and verified (.build/Tyrell.app bundles Contents/MacOS/tyrellbar, universal2, signed, Team ZU2882L4HT) but installed on NO host — bar-in-app measured - on all six. redeploy_app_fleet.zsh proposes canary → rdmbair15m5, sha256 bfa249f11e1ab9ba7af145de4f0bc40994664d8759c1c61737307e6b0f8b9e49, cdhash 511e6f69a2390b0420b3b522406e7405898d4036c2cb2a473cbd0a66e81243c9. Executing requires explicit --execute-* flags AND remote sudo -n. Not run — owner-gated.

No secrets are recorded in this changelog.