Fleet changelogs · dev.ecs0.net
rdmbair15m5-changelog-20260823-0347-fleet-state-reporter-and-patience-rule

Fleet state reporter deployed + Patience Budget rule proposed

Deployed an on-demand/hourly agent-state reporter to all six fleet Macs, ran a fleet-wide state probe, and proposed a fleet rule bounding how long any agent waits on its own work.

Host: rdmbair15m5 · Window: 2026-08-23 03:10 → 03:47 EDT · Operator: claude@rdmbair15m5

Delivered

# Item Location
1 probe_host_state.sh v2.1 — read-only state probe ~/scripts/ on all 6 hosts
2 fleet_state_report.zsh — runs probe, delivers to hub ~/scripts/ on all 6 hosts
3 LaunchAgent com.eastcoastscience.fleetstate — hourly, RunAtLoad, Nice 10, LowPriorityIO loaded on all 6
4 Fleet-wide state report bus 20260823-033224-95B80757
5 Patience Budget rule proposal rdmsm4x:~/dev/RULE-agent-patience-budget-20260823.md, bus 20260823-034531-631488F8

Verified: 6/6 hosts delivering to rdmsm4x:~/dev/apps/fleetreg/state/ (10 files, 6 distinct hosts). jdmbair13m5 reached via the rdmsm4x inbound hop per FLEET.md.

Defects I introduced and fixed (all mine)

  1. v1 _t() helper used ${@:2} — misexpands under zsh, fed timeout malformed args. Filled rdmsm4x's log with Try 'timeout --help'. Removed entirely.
  2. v1 walked ~/Library/Mobile Documents unbounded on every hourly run. Minutes per run on hosts with ~17k zero-byte files. Now behind --deep.
  3. Diagnosed the slowness twice by guessing, both wrong. Blamed the iCloud walk, then 7× zsh -lc login shells (measured: 0s). Profiling took 20s and found the real cause: brctl status = 70s (68 containers here, 106 on rdmsm4x). Moved behind --deep. Result: 80s → 11.5s.
  4. First fleet tool probe reported zero CLI tools on 5 of 6 hosts. False — non-login SSH drops Homebrew from PATH (ISS-001, documented, and I walked into it). Re-probed via zsh -lc: claude + agy + codex are on ALL SIX.

Fleet findings

Reversibility

Outstanding