Fleet changelogs · dev.ecs0.net
rdmbair15m5-changelog-20260823-0258-monitor-page-on-dev-dataroo-net

Live agent monitor published to dev.dataroo.net/monitor/

Host: rdmbair15m5 (source) → rdmsm4x (serving) · Session: richh-69 · Time: 2026-08-23 02:55 → 02:58 EDT Ask: Rich — "add ability to monitor the watch.log from dev.dataroo.net"

URLs

Path URL Auth
Tailnet https://rdmsm4x-1.kangaroo-kitefin.ts.net/monitor/ device-authenticated
LAN http://192.168.0.29:8787/monitor/ trusted network
Public https://dev.dataroo.net/monitor/ existing Basic auth

Verified: tailnet HTTP 200, 5,764 B; LAN HTTP 200; public HTTP 401 (auth in front, as designed).

What the page shows

Auto-refreshes every 30 s. Four sections, all sourced from rdmbair15m5:

  1. agy sessions, live — every agy tab with state / quota banner / subagent count, colour-coded (red = quota-blocked, amber = working, green = idle-healthy).
  2. quota deadlines being tracked — retry-after time, time remaining, and whether a retry was sent.
  3. agy-quota-watch/watch.log — last 80 lines. This was the actual request.
  4. devmon alerts — last 25.

Content assertions passed on the served copy: QUOTA BLOCKED ×3, retry after ×4, ttys001 ×3, watch.log ×2.

Why a separate monitor/ subdirectory

The wiki is regenerated by ~/dev/scripts/gen_site.py every 5 minutes. I checked before writing: it contains no rmtree/os.remove/unlink — it only writes index.html, fleet.html and p/*.html, and never deletes. So an independent monitor/ subtree cannot be clobbered by it, and nginx's location / serves subdirectories without config changes. Hand-editing index.html to add a link would not survive, so the page is reached by direct URL.

Security

Mechanism

~/.agent-coordination/publish_monitor.zsh renders a self-contained page (no CDN, no external fonts, light/dark via prefers-color-scheme) and scps it to rdmsm4x:~/dataroo.net/wiki/monitor/index.html. It resolves rdmsm4x through fleet_host.zsh rather than the bare name, because the bare MagicDNS name still points at the dead tailnet twin.

LaunchAgent net.dataroo.monitorpublish, StartInterval 60, RunAtLoad. Confirmed executing under launchd — publish.out recorded a successful publish and publish.err is 0 bytes, last exit code = 0.

Aqua caveat: the live agy-tab section needs AppleScript, so it is only populated when the publisher runs in the GUI session on this host. The log and deadline sections work regardless. If the page ever shows "no agy tabs visible", the publisher ran outside Aqua — the log content is still valid.


POSTSCRIPT — the LaunchAgents were never firing, and neither was devmon

While verifying the publisher's cadence I found it had run once and never again. Chasing that uncovered something worse.

devmon on rdmbair15m5 had been dead for six hours. 2 samples, newest 360 minutes old, while rdmpw3275m had 358. launchctl print reported it LOADED the whole time. My earlier report that "devmon is running on all five hosts" was true of the registration and false of the execution on this host.

Root cause, from launchd's own log:

launchd[1] [gui/501]: pending spawn, domain in on-demand-only mode: net.dataroo.monitorpublish
backgroundtaskmanagementd: copyJobWithLabel for label net.dataroo.monitorpublish created a job

The gui/501 domain is in on-demand-only mode, so StartInterval jobs are registered but never spawned; only XPC/socket-triggered launches run. New LaunchAgents also register with Background Task Management pending user approval on modern macOS.

Two hypotheses were tested and rejected before landing on this:

Fix: one long-lived supervisor instead of repeated spawns. ~/.agent-coordination/monitor_supervisor.zsh — a single process that loops every 60 s and runs the quota watcher, the monitor publisher and a devmon sample itself. An already-running process is not subject to on-demand-only mode.

Verified working:

devmon samples   2 (stuck 6h)  ->  3  ->  4
page 'generated' 02:56:15      ->  03:06:15

The LaunchAgents are left registered and are harmless — every task they run is idempotent (the quota watcher is guarded by its state file), so if the domain ever leaves on-demand-only mode, double execution does no damage.

Owner action for the durable fix: approve these in System Settings → General → Login Items & Extensions (they appear as background items), so launchd will honour their intervals without a supervisor process.

Lesson worth generalising: launchctl print saying loaded / state = active is NOT evidence a scheduled job is running. Check runs =, and check that the job's output is advancing. Every claim in this session about a LaunchAgent "running" should have been verified that way.