rdmsm4x changelog 20260925-1021 — fleet_audit boot-age guard
Stopped the fleet audit from calling a freshly-booted host WEDGED,
and fixed a greedy kern.boottime parse that had made the
first version of that guard completely inert.
- Host: rdmsm4x · When: 2026-09-25 10:21:05 EDT · Agent: claude@rdmsm4x (session_01V624TifjnsJ7tsa2kxhmWi)
- Budget: Anthropic weekly quota reset; no hold. Worked inline, no subagents.
What changed
| File | Change |
|---|---|
~/dev/scripts/fleet_audit.zsh |
boot-age guard + anchored boottime parse — commit
6a802db |
~/dev/scripts/SESSION-STATE.md |
session 5b4da566 section (append-only) |
fleet_audit.zsh.bak-claude-20260925-1012 |
backup of grok's prior version |
Pushed to backup and fleet.
Why
The 2026-09-24 22:59:01 audit graded rdmpw3275m
WEDGED — "2 launchd job(s) pended, monitoring NOT running", 5
agents vs 16 claimed. That host booted at 22:58:34 — 27 seconds
earlier. agentstatus/agentheal had not reached their first
interval and no agent session had started, so both readings were true
and the verdict was wrong. Re-measured 23:11:
runs=2 / last exit code=0 on both jobs, herdr pid 3793,
pgrep -x claude = 11. Its 06:52 audit: ok
17/17.
A false WEDGED is worse than no audit — it sends someone chasing a ghost and discounts the next real one. It cost ~20 min and two wrong conclusions of mine before boot timestamps settled it.
The second defect — the guard was inert and looked right
sysctl -n kern.boottime prints
{ sec = 1790305062, usec = 40649 } .... A greedy
.*sec = matches through the u of
usec, capturing MICROSECONDS:
sed -E 's/.*sec = ([0-9]+).*/\1/' -> 40649 WRONG (usec)
sed -E 's/^\{ sec = ([0-9]+).*/\1/' -> 1790305062 right
Elapsed came out ~1.79e9, so no threshold could ever be crossed — every value a plausible integer. Any other fleet script computing uptime that way is silently broken the same way.
Verification — three runs on the real fleet
| threshold | expected | observed |
|---|---|---|
| 1800s (default) | no regression | all six hosts ok, 0 warming rows |
| 999999s | guard fires at all | all six WARMING UP |
| 4000s | selective | only rdmbair15m5 (up 63m) warms; other five
ok |
Elapsed cross-checked against uptime(1): rdmsm4x 40687s
= 678m vs up 11:18.
Not changed, deliberately
The agent counting. pgrep -x claude is
correct; -f returns 43 on rdmpw3275m
because it sweeps in "Claude Helper", "Claude Helper (Plugin)",
"disclaimer" and the probing ssh command line. I suspected
-x first (the documented sshd-session trap) and was
wrong.
Authority
fleet_audit.zsh is grok's file (modified 22:50). Routed
to grok 23:15 as bus 20260924-231510-BE4B1BB5; no reply in
~19h, so decided under SA-03 by claude@rdmsm4x and
recorded. Follow-up bussed as 20260925-101835-84A2FA19.
Undo
cd ~/dev/scripts && git revert 6a802db, or
restore fleet_audit.zsh.bak-claude-20260925-1012.
Outstanding
Nothing for Rich. The one untested path is a real fleet reboot
producing WARMING UP instead of WEDGED — it
cannot be run on demand.