Live agent monitor published to dev.dataroo.net/monitor/
Host: rdmbair15m5 (source) → rdmsm4x (serving) ·
Session: richh-69 · Time:
2026-08-23 02:55 → 02:58 EDT Ask: Rich — "add ability
to monitor the watch.log from dev.dataroo.net"
URLs
| Path | URL | Auth |
|---|---|---|
| Tailnet | https://rdmsm4x-1.kangaroo-kitefin.ts.net/monitor/ |
device-authenticated |
| LAN | http://192.168.0.29:8787/monitor/ |
trusted network |
| Public | https://dev.dataroo.net/monitor/ |
existing Basic auth |
Verified: tailnet HTTP 200, 5,764 B; LAN HTTP 200; public HTTP 401 (auth in front, as designed).
What the page shows
Auto-refreshes every 30 s. Four sections, all sourced from rdmbair15m5:
- agy sessions, live — every agy tab with state / quota banner / subagent count, colour-coded (red = quota-blocked, amber = working, green = idle-healthy).
- quota deadlines being tracked — retry-after time, time remaining, and whether a retry was sent.
agy-quota-watch/watch.log— last 80 lines. This was the actual request.- devmon alerts — last 25.
Content assertions passed on the served copy:
QUOTA BLOCKED ×3, retry after ×4,
ttys001 ×3, watch.log ×2.
Why a separate
monitor/ subdirectory
The wiki is regenerated by ~/dev/scripts/gen_site.py
every 5 minutes. I checked before writing: it contains no
rmtree/os.remove/unlink
— it only writes index.html, fleet.html and
p/*.html, and never deletes. So an independent
monitor/ subtree cannot be clobbered by it, and nginx's
location / serves subdirectories without config changes.
Hand-editing index.html to add a link would not
survive, so the page is reached by direct URL.
Security
- Sits behind the existing auth_proxy Basic auth — no new exposure, no new port, no tunnel change.
- The publisher scrubs token-shaped strings
(
tskey-,ghp_,sk-,Bearer …) before writing, even though these sources should never contain any. Verified zero matches on the served page. - Sources are logs and process state only:
watch.log,.agy_quota_state,devmon/alerts.jsonl.
Mechanism
~/.agent-coordination/publish_monitor.zsh renders a
self-contained page (no CDN, no external fonts, light/dark via
prefers-color-scheme) and scps it to
rdmsm4x:~/dataroo.net/wiki/monitor/index.html. It resolves
rdmsm4x through fleet_host.zsh rather than the bare name,
because the bare MagicDNS name still points at the dead tailnet
twin.
LaunchAgent net.dataroo.monitorpublish,
StartInterval 60, RunAtLoad. Confirmed
executing under launchd — publish.out recorded a successful
publish and publish.err is 0 bytes,
last exit code = 0.
Aqua caveat: the live agy-tab section needs AppleScript, so it is only populated when the publisher runs in the GUI session on this host. The log and deadline sections work regardless. If the page ever shows "no agy tabs visible", the publisher ran outside Aqua — the log content is still valid.
POSTSCRIPT — the LaunchAgents were never firing, and neither was devmon
While verifying the publisher's cadence I found it had run once and never again. Chasing that uncovered something worse.
devmon on rdmbair15m5 had been dead for six hours. 2
samples, newest 360 minutes old, while rdmpw3275m had 358.
launchctl print reported it LOADED the whole
time. My earlier report that "devmon is running on all five hosts" was
true of the registration and false of the execution on
this host.
Root cause, from launchd's own log:
launchd[1] [gui/501]: pending spawn, domain in on-demand-only mode: net.dataroo.monitorpublish
backgroundtaskmanagementd: copyJobWithLabel for label net.dataroo.monitorpublish created a job
The gui/501 domain is in on-demand-only
mode, so StartInterval jobs are registered but
never spawned; only XPC/socket-triggered launches run. New LaunchAgents
also register with Background Task Management pending
user approval on modern macOS.
Two hypotheses were tested and rejected before landing on this:
ProcessType: Background+LowPriorityIOthrottling under memory pressure — plausible given pressure level 2 and 13 GB of swap, but removing those keys changed nothing.- The machine cannot spawn processes — disproved directly:
launchd spawned
endpointsecuritydandmdworker_sharedseconds earlier, and a manualzsh -cspawn succeeded.
Fix: one long-lived supervisor instead of repeated
spawns.
~/.agent-coordination/monitor_supervisor.zsh — a single
process that loops every 60 s and runs the quota watcher, the monitor
publisher and a devmon sample itself. An already-running process is not
subject to on-demand-only mode.
Verified working:
devmon samples 2 (stuck 6h) -> 3 -> 4
page 'generated' 02:56:15 -> 03:06:15
The LaunchAgents are left registered and are harmless — every task they run is idempotent (the quota watcher is guarded by its state file), so if the domain ever leaves on-demand-only mode, double execution does no damage.
Owner action for the durable fix: approve these in System Settings → General → Login Items & Extensions (they appear as background items), so launchd will honour their intervals without a supervisor process.
Lesson worth generalising:
launchctl print saying loaded /
state = active is NOT evidence a scheduled job is running.
Check runs =, and check that the job's output is
advancing. Every claim in this session about a LaunchAgent "running"
should have been verified that way.