Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260830-1227-watchdog-panic-diagnosis-and-fleet-storm-guard

rdmsm4x-changelog-20260830-1227-watchdog-panic-diagnosis-and-fleet-storm-guard

2026-08-30 12:27:54 EDT — Diagnosed the rdmsm4x reboot and deployed an automatic fleet-wide guard for its cause.

Why it rebooted

Kernel panic 2026-08-30 01:51:46 EDT, back up 01:53:16.

panic(cpu 0 caller 0xfffffe0049cc2960): watchdog timeout:
no checkins from watchdogd in 93 seconds

Kexts in backtrace: AppleARMWatchdogTimer, AppleInterruptControllerV3. Report: /Library/Logs/DiagnosticReports/panic-base+socd-2026-08-30-120158.000.panic Mac16,9, macOS 27.0 build 26A5421a (beta).

Not a driver fault. watchdogd missed its check-in for 93 s because the system was resource-starved. logd stopped recording at 01:50:07 — ~99 s before the panic — so watchdogd was not the only thing denied a core. Contributing state in the pre-panic window:

Correction to an earlier reading in this session

The 12:01:57 event today was not a second incident. /dev/console is stamped 12:02:48 and launchd (pid 1) still dates to 01:53:16: the Mac ran with no Aqua session for 10h09m after the panic reboot, and 12:01 was simply the login. DumpPanic ran then because login is when a deferred panic report is surfaced. The load-average spike of 801 at 12:03 was the login thundering herd. Load average is misleading on this host generally: load1 sits 60–100 on 16 cores while CPU is 71–97% idle. Cross-checked with dev-bf (claude@rdmsm4x), who measured /dev/console independently.

What was deployed

com.eastcoastscience.stormguard — StartInterval 300 s, RunAtLoad, Nice 5.

Design points that matter:

Fleet state

Host Status
rdmsm4x INSTALLED, verified sha 58c58cd39583, runs=1
rdmbair13m5 INSTALLED, sha MATCH
rdmpw3265m INSTALLED, sha MATCH
rdmpw3275m INSTALLED, sha MATCH
jdmbair13m5 INSTALLED, sha MATCH
rdmbair15m5 PENDING — asleep (Bonjour sleep proxy answers ping; Tailscale offline 24m). No MAC recorded so no WoL. Lead retries hourly via .pending_hosts and self-converges.

Verification performed

Fleet traps avoided (per mac-fleet-host-onboarding)

Undo

launchctl bootout gui/$(id -u)/com.eastcoastscience.stormguard
rm ~/Library/LaunchAgents/com.eastcoastscience.stormguard.plist
rm ~/.agent-coordination/storm_guard.zsh
# repeat per host; logs under ~/.agent-coordination/stormguard/ can be kept

Owner actions outstanding


ADDENDUM 2026-08-30 12:34:24 EDT — the likely panic mechanism, found after the guard was deployed

The Time Machine backup on rdmsm4x is stalled, and it is the best explanation for the watchdog starvation.

Why this matters: threads blocked on a stuck network filesystem count toward macOS load average. That is why load1 reads 60-100 while the CPU sits 71-97% idle — the load was never CPU demand. A kernel thread wedged behind stalled SMB I/O is the most plausible route to watchdogd missing its check-in for 93 seconds.

Every one of those four processes is on the guard's NEVER list, which is correct — killing backupd or diskimagesiod mid-backup is not an automation's call. The guard was extended to detect the stall specifically (progress < 1 MB between passes while Running=1) and escalate it, report-only.

This one needs Rich. It is a NAS/SMB problem: check unaspro818a, the share, and the sparsebundle. Until it is fixed the host stays exposed to the same starvation.

Correction recorded

An earlier ISSUES.md entry claimed rdmbair15m5 had no MAC in fleet-macs.json. That was wrong — my probe walked only the top level of the JSON; the entry exists at hosts.rdmbair15m5 (en0 fc:b2:14:41:58:bd). Corrected in place. Per dev-bf, WoL fails for a different reason (the address answering on the wire is locally-administered — sleep proxy or randomized Wi-Fi MAC) and that host also refuses the fleet ssh key, so it needs Rich at the keyboard regardless. fleet-macs.json needs no edit.

Final fleet state 2026-08-30 12:34:24 EDT

Host Label SHA ecb6cd3553a2
rdmsm4x LIVE MATCH
rdmbair13m5 LIVE MATCH
rdmpw3265m LIVE MATCH
rdmpw3275m LIVE MATCH
jdmbair13m5 LIVE MATCH (first probe timed out; re-verified)
rdmbair15m5 pending asleep; lead self-converges hourly

ADDENDUM 2 2026-08-30 12:43:22 EDT — two defects found in the guard's own escalation path

dev-bf named a pattern while correcting their own backup finding: a presence check reported as a delivery check. tmutil destinationinfo proves a destination is configured; tmutil latestbackup proves it works. Applying that test to this script found the same error twice, both in the one path whose entire job is telling a human.

1. Escalation reported success without checking delivery. The once-per-day rate-limit key was written BEFORE any send was attempted, both channels ran under >/dev/null 2>&1 with their exit codes discarded, and say "ESCALATED" logged unconditionally. A failed send during a real storm was indistinguishable from a successful one and suppressed every retry for 24 hours. Fixed: each channel is checked by exit code, the key is burned only if something got through, an undelivered escalation logs loudly and retries next pass, and the body is written to ~/.agent-coordination/stormguard/escalation-<ts>.txt first and unconditionally so evidence outlives a failed attempt.

2. The ticket channel had never worked. ticket open takes --title/--type/--desc; the script passed a positional title and --body, which argparse rejected every time. Invisible because of defect 1. Verified working form now in place.

3. Ordering bug: the stalled-backup escalation could never fire. The Time Machine check ran AFTER the escalation test that reads TMSTALL, so ${TMSTALL:-0} silently evaluated false on every pass. The :-0 default masked it — no error under set -u, just a permanently dead branch, in the guard's single most important detection. Block moved ahead of the escalation test.

Verified end to end against the live stalled backup:

12:41:18  ESCALATED [tm_stalled=1] delivered=bus  failed=ticket   <- caught defect 2
12:42:01  ESCALATED [tm_stalled=1] delivered=bus,ticket failed=none

Real escalation raised: ISSUE-20260830-38 (incident, high, fleet-maintenance). Selftest ticket ISSUE-20260830-37 opened and resolved.

Redeployed: sha a28207a40535, verified on rdmsm4x, rdmbair13m5, rdmpw3265m, rdmpw3275m, jdmbair13m5 (sha match + label loaded). rdmbair15m5 still pending, self-converges.