Fleet changelogs · dev.ecs0.net
rdmbair15m5-changelog-20260823-1826-agent-fleet-self-healing-and-4h-stability-review

rdmbair15m5 - changelog - agent-fleet-self-healing-and-4h-stability-review - claude - fleet-automation - agent-coordination - 20260823-1825

Session: 2026-08-22 21:31 EDT -> 2026-08-23 18:26 EDT (~21 h) Ran on: rdmbair15m5 · Changed state on: all six fleet hosts, canonical work on rdmsm4x

Built an automated agent-fleet management layer: heartbeats, automatic bus sync, self-healing, independent off-host auditing with remote repair, a status board, and a 4-hourly stability review graded on delivery ratio. Rich's stated goal was to stop managing agents by hand — "it is enough just directing work, vs. managing and checking all agent work."


What now runs unattended

Job Cadence Scope Purpose
com.eastcoastscience.agentstatus 5 min all 6 agent heartbeat + automatic bus sync
com.eastcoastscience.agentheal 10 min all 5 self-heal wedged launchd; escalate what it cannot
com.eastcoastscience.offhostwatch 10 min rdmbair15m5 watch rdmsm4x from outside
com.eastcoastscience.fleetaudit 15 min rdmsm4x independent verification + remote healing
com.eastcoastscience.stabilityreview 4 h rdmsm4x delivery-ratio grading

All carry RunAtLoad=true. Scripts: ~/.agent-coordination/{agent_status,agent_heal}.zsh, ~/dev/_ops/fleetwatch/offhost_watch.zsh, rdmsm4x:~/dev/scripts/{fleet_audit,fleet_stability_review}.zsh.

The two gaps that existed before this

  1. The bus never synced on its own. An inventory of every LaunchAgent on all five hosts found no job anywhere running agent_msg.zsh sync — including on rdmsm4x, the hub. It ran only when an agent remembered to type it.
  2. Nothing detected a running-but-silent agent. No liveness signal existed at all. rdmpw3275m sent its first-ever bus message at 03:34:53 once this shipped.

Agent status board

rdmsm4x:~/dev/scripts/gen_agent_board.py -> ~/dataroo.net/wiki/agents.html, wired into publish_site.zsh and the gen_site.py nav. Live at https://dev.dataroo.net/agents.html.

dev.dataroo.net — two fixes

Config normalization

Fixed 5 contradictions; verified 5 others as already-correct and deliberately did not churn them. Record: rdmsm4x:~/dev/fleet/NORMALIZATION-20260823.md. Notably FLEET.md pointed at a project path that exists on no host and contradicted CLAUDE.md; the fleet-notes-publish skill named two scripts that do not exist and described the retired per-host Notes folders. Added rule 28 to AGENT_COORDINATION.md (stay on task, route unrelated work to claude@rdmsm4x), propagated fleet-wide. Created 11 templates at rdmsm4x:~/dev/lib/templates/.

First stability review — DEGRADED, fleet delivery 75%

rdmpw3265m 100%   rdmpw3275m 100%   rdmbair13m5 95%   jdmbair13m5 93%
rdmsm4x     60%   rdmbair15m5  2%  <- FAILING, 4 wedge events

Second review, 17:22, four hours later, unattended — DEGRADED, fleet 72%

rdmpw3265m  99%   rdmpw3275m  99%   jdmbair13m5 93%   rdmbair13m5 unreachable
rdmsm4x     59%   rdmbair15m5 11%   <- still FAILING

rdmbair15m5 improved 2% -> 11% over five hours WITHOUT a reboot: 19 of ~164 expected runs. The remote heal from rdmsm4x keeps rescuing it; launchd keeps refusing to run it unaided. That is the distinction the metric exists to make — the host is being kept alive FROM OUTSIDE, not recovering. Its three jobs read ok with runs 23/22/25, which is exactly the misleading snapshot the delivery ratio sees through. The 4 h cycle itself is proven: it ran on time, unattended, twice.

The finding that changed the design

Self-healing cannot fix a broken launchd. agentheal was itself wedged by the exact condition it repairs — launchd is what spawns the healer. Verified: at 12:50 it healed two jobs; by 13:17 all three including itself were pended. Only remote healing from rdmsm4x recovered it. Cadence raised 3600s -> 900s accordingly; a full audit costs 7 seconds.

Method note

Eighteen defects across five agents this session, every one the same shape: a proxy standing in for the thing. mtime for write-time · transcript text for execution · exit status for data health · tail's status for the command's · pgrep -f for a running program · a monitor's log for the monitor running · state = not running for a job that is stuck pending · host uptime for how long a job should have been running. Measure the thing itself. Four-plus were caught only because another agent checked the first one's work.

Outstanding for Rich

  1. Reboot — in progress. rdmbair15m5's launchd is wedged and self-repair provably cannot hold it.
  2. LogTTY authoritative tree · DevSort commit authority · the unidentified LogTTY copier — all parked.
  3. devsite 60% / fleethealth 77% — likely a benign cadence overrun, left with claude@rdmsm4x.
  4. Distribution certificates: zero on the fleet. Notarization and App Store blocked.

Verification

All job states read via launchctl print (runs, last exit code, pended nondemand spawn) rather than from output files — a heartbeat file is exactly what a hand-run also produces.

Secrets

None written. Scripts read credentials at runtime from ~/.secrets/global.env.