rdmsm4x-changelog-20260830-1227-watchdog-panic-diagnosis-and-fleet-storm-guard
2026-08-30 12:27:54 EDT — Diagnosed the rdmsm4x reboot and deployed an automatic fleet-wide guard for its cause.
Why it rebooted
Kernel panic 2026-08-30 01:51:46 EDT, back up 01:53:16.
panic(cpu 0 caller 0xfffffe0049cc2960): watchdog timeout:
no checkins from watchdogd in 93 seconds
Kexts in backtrace: AppleARMWatchdogTimer, AppleInterruptControllerV3. Report: /Library/Logs/DiagnosticReports/panic-base+socd-2026-08-30-120158.000.panic Mac16,9, macOS 27.0 build 26A5421a (beta).
Not a driver fault. watchdogd missed its check-in for 93 s because the system was resource-starved. logd stopped recording at 01:50:07 — ~99 s before the panic — so watchdogd was not the only thing denied a core. Contributing state in the pre-panic window:
- JetsamEvent 00:33:17 — compressor 2.9 GB, 40.1M compressions, 25.2M decompressions.
- biomesyncd CPU-resource violation covering 01:46:59–01:49:20, ending 2m26s before the panic.
- 1393 processes at 00:33.
Correction to an earlier reading in this session
The 12:01:57 event today was not a second incident.
/dev/console is stamped 12:02:48 and launchd (pid 1) still
dates to 01:53:16: the Mac ran with no Aqua session for 10h09m after the
panic reboot, and 12:01 was simply the login. DumpPanic ran then because
login is when a deferred panic report is surfaced. The load-average
spike of 801 at 12:03 was the login thundering herd. Load average is
misleading on this host generally: load1 sits 60–100 on 16 cores while
CPU is 71–97% idle. Cross-checked with dev-bf (claude@rdmsm4x), who
measured /dev/console independently.
What was deployed
com.eastcoastscience.stormguard — StartInterval 300 s,
RunAtLoad, Nice 5.
- Canonical:
~/dev/fleet/maintenance/scripts/host_storm_guard.zsh(v1.0) - Per-host:
~/.agent-coordination/storm_guard.zsh - Log/state:
~/.agent-coordination/stormguard/(stormguard.log, history.tsv, .flagged, .pending_hosts) - Agent:
~/Library/LaunchAgents/com.eastcoastscience.stormguard.plist
Design points that matter:
- Measures true CPU rate, sampling accumulated
CPU-time twice over 5 s.
ps %cpuis a decaying average since process start; an allowlist acting on it kills healthy processes all day. - Trigger is measured busy%, never load average, for the reason above.
- Two-strike policy: >=150% of one core acts at once; >=60% must also have been flagged on the previous pass. Legitimate first-pass work gets to finish.
- Kills only an explicit allowlist of stateless Apple sync daemons that respawn clean.
- Never kills mds_stores/corespotlightd (would force a full reindex), backupd, diskimagesiod, bird, fileproviderd, FPCKService, or anything holding user work (claude, codex, agy, builds).
- Memory pressure and process population are reported, never remediated — the processes holding that memory are Rich's work.
- Orphan sweep confirms an orphan by absence from
launchctl list, not by ppid==1 alone (every launchd job has ppid 1). - Escalates via agent_msg.zsh to claude@rdmsm4x and
ticket open incident, one per host per day.
Fleet state
| Host | Status |
|---|---|
| rdmsm4x | INSTALLED, verified sha 58c58cd39583, runs=1 |
| rdmbair13m5 | INSTALLED, sha MATCH |
| rdmpw3265m | INSTALLED, sha MATCH |
| rdmpw3275m | INSTALLED, sha MATCH |
| jdmbair13m5 | INSTALLED, sha MATCH |
| rdmbair15m5 | PENDING — asleep (Bonjour sleep proxy answers ping; Tailscale offline 24m). No MAC recorded so no WoL. Lead retries hourly via .pending_hosts and self-converges. |
Verification performed
- Syntax check, dry run on live system.
- Kill path proven end to end with a compiled CPU
burner named
postersyncd: pass 1 flagged (strike 1, no kill), pass 2 killed at 83.3% via SIGTERM, recorded in history.tsv. - launchd delivery confirmed by run count, not by
state: runs=1 after RunAtLoad on all 5 hosts. - SHA parity verified per host against the canonical file.
Fleet traps avoided (per mac-fleet-host-onboarding)
$HOSTis a zsh built-in equal to the local hostname — renamed toTHIS_HOST, with a hard abort if the ComputerName probe returns empty. This exact trap has previously run a full fleet sequence against the lead host reporting exit 0 throughout.ssh -neverywhere in loops (ssh drains stdin, so a loop runs once and still reports success).- Host pinned by SSH host-key fingerprint with a loopback refusal; ComputerName alone proves nothing.
- Atomic install via
.new+mv— zsh reads scripts lazily by byte offset, so overwriting one mid-run yields a parse error pointing at valid code.
Undo
launchctl bootout gui/$(id -u)/com.eastcoastscience.stormguard
rm ~/Library/LaunchAgents/com.eastcoastscience.stormguard.plist
rm ~/.agent-coordination/storm_guard.zsh
# repeat per host; logs under ~/.agent-coordination/stormguard/ can be keptOwner actions outstanding
- rdmbair15m5 will self-converge on wake; confirm
with
storm_guard.zsh --statusthere later. - Unclassified hot processes seen on rdmsm4x and left untouched pending a decision: 1Password (77–217%), FPCKService (65–74%), log. Classify into KILLABLE or NEVER in the script.
- Process population 1256–1267 on rdmsm4x is near the 1393 that preceded the panic. The orphan sweep only reclaims MCP helpers older than 60 min; the rest is live agent sprawl.
- No secrets are written by this job. No credentials appear in any artifact.
ADDENDUM 2026-08-30 12:34:24 EDT — the likely panic mechanism, found after the guard was deployed
The Time Machine backup on rdmsm4x is stalled, and it is the best explanation for the watchdog starvation.
tmutil status: Running=1, phase Copying, destinationsmb://rich@unaspro818a/timemachine.- Progress: 76 files of 20,259,556; 13 MB of 3.9 TB. Measured 0 bytes in 15 s and 0 bytes across a full 5-minute guard pass. One sample showed 614 B/s.
backupd,diskimagesiod,mds_storesandFPCKServiceare all in uninterruptible wait.
Why this matters: threads blocked on a stuck network filesystem count toward macOS load average. That is why load1 reads 60-100 while the CPU sits 71-97% idle — the load was never CPU demand. A kernel thread wedged behind stalled SMB I/O is the most plausible route to watchdogd missing its check-in for 93 seconds.
Every one of those four processes is on the guard's NEVER list, which is correct — killing backupd or diskimagesiod mid-backup is not an automation's call. The guard was extended to detect the stall specifically (progress < 1 MB between passes while Running=1) and escalate it, report-only.
This one needs Rich. It is a NAS/SMB problem: check unaspro818a, the share, and the sparsebundle. Until it is fixed the host stays exposed to the same starvation.
Correction recorded
An earlier ISSUES.md entry claimed rdmbair15m5 had no MAC in
fleet-macs.json. That was wrong — my probe walked only the top level of
the JSON; the entry exists at hosts.rdmbair15m5 (en0
fc:b2:14:41:58:bd). Corrected in place. Per dev-bf, WoL fails for a
different reason (the address answering on the wire is
locally-administered — sleep proxy or randomized Wi-Fi MAC) and that
host also refuses the fleet ssh key, so it needs Rich at the keyboard
regardless. fleet-macs.json needs no edit.
Final fleet state 2026-08-30 12:34:24 EDT
| Host | Label | SHA ecb6cd3553a2 |
|---|---|---|
| rdmsm4x | LIVE | MATCH |
| rdmbair13m5 | LIVE | MATCH |
| rdmpw3265m | LIVE | MATCH |
| rdmpw3275m | LIVE | MATCH |
| jdmbair13m5 | LIVE | MATCH (first probe timed out; re-verified) |
| rdmbair15m5 | pending | asleep; lead self-converges hourly |
ADDENDUM 2 2026-08-30 12:43:22 EDT — two defects found in the guard's own escalation path
dev-bf named a pattern while correcting their own backup finding:
a presence check reported as a delivery check.
tmutil destinationinfo proves a destination is configured;
tmutil latestbackup proves it works. Applying that test to
this script found the same error twice, both in the one path whose
entire job is telling a human.
1. Escalation reported success without checking
delivery. The once-per-day rate-limit key was written BEFORE
any send was attempted, both channels ran under
>/dev/null 2>&1 with their exit codes discarded,
and say "ESCALATED" logged unconditionally. A failed send
during a real storm was indistinguishable from a successful one and
suppressed every retry for 24 hours. Fixed: each channel is checked by
exit code, the key is burned only if something got through, an
undelivered escalation logs loudly and retries next pass, and the body
is written to
~/.agent-coordination/stormguard/escalation-<ts>.txt
first and unconditionally so evidence outlives a failed attempt.
2. The ticket channel had never worked.
ticket open takes --title/--type/--desc; the
script passed a positional title and --body, which argparse
rejected every time. Invisible because of defect 1. Verified working
form now in place.
3. Ordering bug: the stalled-backup escalation could never
fire. The Time Machine check ran AFTER the escalation test that
reads TMSTALL, so ${TMSTALL:-0} silently
evaluated false on every pass. The :-0 default masked it —
no error under set -u, just a permanently dead branch, in
the guard's single most important detection. Block moved ahead of the
escalation test.
Verified end to end against the live stalled backup:
12:41:18 ESCALATED [tm_stalled=1] delivered=bus failed=ticket <- caught defect 2
12:42:01 ESCALATED [tm_stalled=1] delivered=bus,ticket failed=none
Real escalation raised: ISSUE-20260830-38 (incident, high, fleet-maintenance). Selftest ticket ISSUE-20260830-37 opened and resolved.
Redeployed: sha a28207a40535, verified on rdmsm4x, rdmbair13m5, rdmpw3265m, rdmpw3275m, jdmbair13m5 (sha match + label loaded). rdmbair15m5 still pending, self-converges.