Fleet changelogs · dev.ecs0.net
rdmbair15m5-changelog-20260823-1120-monitoring-fixes-applied-overnight-resume-verified

Monitoring fixes applied; overnight quota resume verified; one wedged session left

Host: rdmbair15m5 · Session: Claude Code richh-69 · Time: 2026-08-23 11:17 → 11:20 EDT

Correction issued first

I told Rich to approve the LaunchAgents in System Settings → Login Items & Extensions. That was wrong. sfltool dumpbtm shows all four already registered as Disposition: [enabled, allowed, notified]. Nothing was pending approval and there was nothing for him to click. net.dataroo.agyquotawatch even records Last Use: 2026-08-23 02:44:48, so the gui/501 domain is only intermittently in on-demand-only mode, not permanently.

The overnight automation worked

watch.log shows the supervisor tracking ttys000's countdown continuously from 04:16 to 05:40 and then acting exactly on the deadline:

05:42:42  /dev/ttys000 RESUME ATTEMPTED (was Resets in 2m9s)
05:53:26  /dev/ttys001 RESUME ATTEMPTED (was Resets in 2h49m0s)

Result at 11:18: no quota banner and no Error ID on any agy tab. ttys001 and ttys007 are working. The supervisor ticked to 420 (7 h of 60 s cycles) without intervention.

Note the earlier apparent "8-hour sleep" was my own session being idle between turns — the machine never slept and the supervisor never stopped.

Fixes applied

Fix Result
Supervisor made reboot-durable LaunchAgent net.dataroo.monitorsupervisor with KeepAlive + RunAtLoad
Four StartInterval agents retired devmon, devmon.fleet, agyquotawatch, monitorpublish unloaded, plists renamed .retired-20260823
Duplicate supervisor removed one process, pid 47005, launchd-managed
Stale quota entries cleared state file empty; re-arms on the next real block

Key discovery: KeepAlive spawns where StartInterval does not. The supervisor agent started immediately under launchd (runs=1) while the interval agents sat at runs=0 with pending spawn, domain in on-demand-only mode. That is the durable pattern for this host — one long-lived KeepAlive process that does its own timing, not many interval jobs.

Everything verified moving

devmon samples   2 (stuck 6h)  ->  449
monitor page     refreshed 11:18:59, 16s old at check
quota tracked    0 tabs blocked
memory           swap 13.1 GB -> 5.29 GB, free 123 MB, pressure still 2

Outstanding — needs Rich

agy ttys000 is wedged and cannot be fixed from outside. All 21 subagents report the identical elapsed time (10h7m35s), which means none is progressing — they stalled when the session hit quota at ~02:40 and did not recover even though the quota reset at 05:42 and the session resumed. The main loop now refuses input: a cleanup directive was delivered and Return pressed twice, and the tab stayed idle at 21 subagents.

I did not restart it. My own fleet rule, written into the standing authorization block this morning, says killing a session holding unsaved context requires asking first — and here I cannot even checkpoint it, because it will not accept the prompt.

Recovery material is captured: ~/.agent-coordination/quota-hold-20260823-0242/ (ttys000-transcript.txt at 02:42 and ttys000-transcript-1120.txt at 11:20, plus a README naming the task and deadline).

Also still open and not mine to do: the five dead Tailscale twin nodes (needs the admin console or a TAILSCALE_API_KEY), and rdmsm4x defaultMode=auto (owned by that host's live session).