Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260925-1021-fleet-audit-boot-age-guard

rdmsm4x changelog 20260925-1021 — fleet_audit boot-age guard

Stopped the fleet audit from calling a freshly-booted host WEDGED, and fixed a greedy kern.boottime parse that had made the first version of that guard completely inert.

What changed

File Change
~/dev/scripts/fleet_audit.zsh boot-age guard + anchored boottime parse — commit 6a802db
~/dev/scripts/SESSION-STATE.md session 5b4da566 section (append-only)
fleet_audit.zsh.bak-claude-20260925-1012 backup of grok's prior version

Pushed to backup and fleet.

Why

The 2026-09-24 22:59:01 audit graded rdmpw3275m WEDGED — "2 launchd job(s) pended, monitoring NOT running", 5 agents vs 16 claimed. That host booted at 22:58:34 — 27 seconds earlier. agentstatus/agentheal had not reached their first interval and no agent session had started, so both readings were true and the verdict was wrong. Re-measured 23:11: runs=2 / last exit code=0 on both jobs, herdr pid 3793, pgrep -x claude = 11. Its 06:52 audit: ok 17/17.

A false WEDGED is worse than no audit — it sends someone chasing a ghost and discounts the next real one. It cost ~20 min and two wrong conclusions of mine before boot timestamps settled it.

The second defect — the guard was inert and looked right

sysctl -n kern.boottime prints { sec = 1790305062, usec = 40649 } .... A greedy .*sec = matches through the u of usec, capturing MICROSECONDS:

sed -E 's/.*sec = ([0-9]+).*/\1/'    -> 40649        WRONG (usec)
sed -E 's/^\{ sec = ([0-9]+).*/\1/'  -> 1790305062   right

Elapsed came out ~1.79e9, so no threshold could ever be crossed — every value a plausible integer. Any other fleet script computing uptime that way is silently broken the same way.

Verification — three runs on the real fleet

threshold expected observed
1800s (default) no regression all six hosts ok, 0 warming rows
999999s guard fires at all all six WARMING UP
4000s selective only rdmbair15m5 (up 63m) warms; other five ok

Elapsed cross-checked against uptime(1): rdmsm4x 40687s = 678m vs up 11:18.

Not changed, deliberately

The agent counting. pgrep -x claude is correct; -f returns 43 on rdmpw3275m because it sweeps in "Claude Helper", "Claude Helper (Plugin)", "disclaimer" and the probing ssh command line. I suspected -x first (the documented sshd-session trap) and was wrong.

Authority

fleet_audit.zsh is grok's file (modified 22:50). Routed to grok 23:15 as bus 20260924-231510-BE4B1BB5; no reply in ~19h, so decided under SA-03 by claude@rdmsm4x and recorded. Follow-up bussed as 20260925-101835-84A2FA19.

Undo

cd ~/dev/scripts && git revert 6a802db, or restore fleet_audit.zsh.bak-claude-20260925-1012.

Outstanding

Nothing for Rich. The one untested path is a real fleet reboot producing WARMING UP instead of WEDGED — it cannot be run on demand.