Fleet directive relayed, 595 false alerts fixed, rooDB migration approved
Host: rdmbair15m5 · Session:
richh-69 · Time: 2026-08-23 11:22 → 11:37
EDT
1. Rich's directive relayed — the intended recipient no longer exists
"tell rdmsm4x to check in with all agents, get state of projects and have it resume/resolve fix issues and get work resumed across the fleet"
rdmsm4x-ticklish-rocket, the session that had been
coordinating rdmsm4x (and owned the LogTTY merge), is gone from
ListAgents. Rather than send into a void, the directive went
out four ways:
SendMessage→dev_(idle, Remote Control)SendMessage→rdmsm4x-validated-cupcake(queued; it is blocked, see below)- durable file →
rdmsm4x:~/dev/_ops/FLEET-RESUME-DIRECTIVE-20260823.md - bus broadcast → all five hosts, topic
fleet-resume-directive-20260823
The brief carries measured state so the recipient does not re-derive it.
2. Measured fleet state, 11:22 EDT
| host | agy | claude | codex | devmon | alerts | open requests |
|---|---|---|---|---|---|---|
| rdmbair15m5 | 6 | 5 | 3 | 460 | 1549 | 9 |
| rdmsm4x | 2 | 5 | 3 | 522 | 8 | 5 |
| rdmbair13m5 | 0 | 2 | 2 | 857 | 595 | 0 |
| rdmpw3265m | 0 | 0 | 1 | 862 | 3 | 0 |
| rdmpw3275m | 0 | 0 | 2 | 863 | 0 | 0 |
rdmsm4x dirty: LogTTY:61, RTTy:20, ScanRoo:11,
rdDB:10, plus 7 more. bookmarkroo and rooDB
now exist there — the renames are underway.
Two peers stuck ~9 hours in
requires_action:
rdmsm4x-validated-cupcake and
jdmbair13m5-tender-whale. Blocked on a permission prompt,
will not self-recover.
3. The 595 alerts were my bug, not an incident — fixed
rdmbair13m5's 595 alerts over 14 hours were all
swap 79–87% used, firing every minute, while kernel
pressure read 1 (normal) and the top consumers were
corespotlightd (1.10 GB), Finder (1.08 GB) and
Notes (0.60 GB). No dev app, no agy, no incident.
macOS sizes the swap file dynamically, so a small swap sits near-full as a matter of course. Alerting on swap percentage alone is the same class of mistake as the free-RAM rule I fixed yesterday: alarming on a proxy instead of the kernel's own verdict.
Fix: swap and compressor alerts now require
kern.memorystatus_vm_pressure_level >= 2. Deployed to
all four peer hosts. Verified both directions:
rdmbair13m5 (pressure 1) 0 new alerts <- 595-per-14h noise gone
rdmbair15m5 1 new alert <- and it is real: agy(67715) rss 4980 MB
A monitor that cries wolf gets ignored, which is worse than no monitor.
4. rooDB migration plan — reviewed and APPROVED with two conditions
agy delivered roodb-migration-plan.md at 11:27. I
verified its claims rather than accepting them:
| Claim | Result |
|---|---|
| Purpose Filter implemented | ✅ enum FilePurpose at
Sources/RDDBCore/Models.swift:73, 22 refs, 32 schema
refs |
| Tests green | ✅ 163 tests, 0 failures, reproduced here |
| Now a git repo | ✅ b778c07 — was not a repo at 22:38 yesterday |
| os_log instrumentation | ⚠️ overstated |
Condition 1: the plan lists the OSLog subsystem as a
rename surface, implying instrumentation is done. Measured:
import OSLog 2, Logger( 8 — but 84 of
the original 93 print() calls remain. Only 9 were
converted. Renaming a subsystem that carries ~10% of the app's output
means doing the rename twice. Finish the conversion first.
Condition 2: do not execute on rdmbair15m5 while it is at pressure 2 with ~123 MB free and a wedged 4.98 GB agy session. Moving a 507 MB SQLite store and re-registering a launchd daemon in that state invites corruption. Run it on rdmsm4x, or wait for pressure to return to 1.
Everything else approved: the rename surface matrix is complete, and
the DB procedure (checkpoint → integrity_check → atomic
copy → re-verify → retain original) is correct.
Review filed to requests-to-claude/, broadcast on the
bus, and delivered into ttys001/003/004.
Still outstanding for Rich
- agy
ttys000wedged: 21 subagents frozen at identical elapsed, main loop refusing input, now 4,980 MB and up 1 d 8 h. Restart needs his word; transcripts captured. - Five dead Tailscale twin nodes — admin console or a
TAILSCALE_API_KEY. - Two peers stuck in
requires_actionfor 9 hours.