rdmsm4x-changelog-20260909-1300-control-plane-monitoring-only-stall-recovery
Fleet execution stall (since ~2026-09-04) unblocked: control-plane automation set to monitoring only, cloud routines off, hub synced with GitHub mirrors. Written 2026-09-09 13:00:12 EDT by claude@rdmsm4x (Fable 5.1) session_01L5FqUPjVWGsMbP4UhAqAa6.
DEC-20260909-01 — Fleet control plane → MONITORING ONLY; work local, no cloud, no containers (2026-09-09 13:00:12 EDT · rdmsm4x)
Rich, live, 2026-09-09 12:4x EDT: "we are seemingly stuck on execution… if stuck on Tyrell, temporarily unhook it and put it back to monitoring mode vs. control plane… un-do containers… we should be working locally… things have seemed stalled since right before Labor Day." Lead: claude@rdmsm4x session_01L5FqUPjVWGsMbP4UhAqAa6 (Fable 5.1, oversight tier).
Root causes found (all installed 2026-09-04/05, i.e. right before Labor Day):
fleet-load-guard(90 s) broadcast a fleet-wide SLOWDOWN directive every ~30–60 min and a recovery notice in between — flip-flop all morning (12:10, 12:35 today; 10:11, 11:28 earlier). Every agent that read the bus was told to stagger dispatches 45–60 s and cut parallelism, repeatedly.usage_statusadvice = conserve (Anthropic weekly 72 %, resets 2026-09-11 10:00 EDT) — by CLAUDE.md that forbids every fan-out on every agent. OpenAI is at 0 % with 2 500 credits and was unused.- Cloud sessions/routines: 5 idle cloud sessions (Tyrell, ReplicantDB,
ecs0lib, RTTy ×2) and 20 routines (2 enabled) were doing "oversight"
from Ubuntu containers with no Swift toolchain, waiting on a
disconnected hub bridge, and merging PRs to the GitHub mirrors
uncompiled ("MERGE_AND_INTEGRATION rule 3"). Hub
checkouts had no
origin; the merges never reached the hub or thefleetbare store. claude-quota-monitorranssh host "claude agents --json"with a Python timeout that kills only the local ssh — 121 orphaned probes on rdmbair13m5 (one every 6 min).stormguard/fleet-realtime-guardkill processes;resume-dispatcher(RunAtLoad) respawned idle sessions into Terminal on every boot. Boot storm at 12:16 today: load1 143 on 16 cores.
- NOT a cause: the OrbStack containers (wiki, Grafana, Loki, cloudflared, certbot, local-llm) are services, not execution sandboxes — left running. Hooks are advisory (never block). No local sandbox/devcontainer config exists.
Actions taken (all reversible):
launchctl bootout+disableon rdmsm4x: fleet-load-guard, fleet-realtime-guard, stormguard, resume-dispatcher, claude-quota-monitor. stormguard also disabled on rdmbair15m5, rdmbair13m5, rdmpw3265m, rdmpw3275m. Plists left in place. Monitoring jobs (tyrelld, tyrellbar, tyrell-usage, agentstatus, fleetstate, fleethealth, fleetaudit, stabilityreview, nightlyaudit, buswatch, …) untouched. configsync does not manage LaunchAgents, so it will not re-enable them.- Killed the 121 orphaned
claude agents --jsonon rdmbair13m5 (before=121 after=0). - Bus broadcast 20260909-125301-D597A727
fleet-slowdown-rescinded-monitoring-only(high). - Cloud routines disabled: trig_01GNiTb5VYTwF6hzPsNvdJYj (replicantDB 6-hourly build ask), trig_01W83o8u75XsL5N1cm3R9Z9a (replicantDB check-in 23:23Z). The 5 idle cloud sessions were NOT woken; their workstreams are taken local.
- Hub brought current with GitHub mirrors (0 open PRs anywhere under
richhdoty): replicantDB ff b3fa0ec→5447787 (+5); lib/app-baseline ff
125d17e→3628c38 (+4); apps/ramdisk merge 3ae37f4→2ba0c3f (+22, 0
overlapping files); sites/cvedb.io 3ca3f5c and
sites/eastcoastscience.com f5b5ba8 (merge -Xours, duplicate generated
PRD); lib/ecs0lib (+12) and net/grafana (+1) pushed to
fleet. All pushed to bothbackupandfleet. - Dispatched ONE Sonnet 5 worker (nice 15) to
swift build/swift testreplicantDB and Tyrell locally and append numbers to each SESSION-STATE.md + ticket ISSUE-20260909-09.
Interpretation of conserve during this
recovery: it means no Claude fan-out beyond one worker at a
time and route bounded implementation to codex exec; it
does NOT mean idle. Rich's explicit unblock instruction today outranks
the gate. Re-evaluate at the weekly reset 2026-09-11 10:00 EDT.
Left alone / needs Rich: apps/macOS-ramDisk-app_eval is a second checkout of the same GitHub repo as apps/ramdisk, diverged with 5 overlapping files (INC-20260905-01 recovery commit) — not merged. Dirty trees not touched: fleet (570), sites/dev.ecs0.net (102), ai/LLM (106), docs (48), issues (15). UNATTENDED marker still present (Rich's).
To re-enable the control plane later:
launchctl enable gui/$(id -u)/<label> && launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/<label>.plist
per job; fix the quota monitor's remote timeout (kill the remote side,
e.g. ssh host 'timeout 5 claude agents --json') before
re-enabling it; re-enable routines at https://claude.ai/code/routines.
Undo
- Re-enable jobs: see the checklist above. Nothing was deleted; plists
untouched; git merges are ordinary merge commits (revert with
git revert -m 1 <sha>).
Follow-through (2026-09-09 13:15:23 EDT)
- Local verification worker (Sonnet 5): Tyrell 9600867 builds, tests 358+80+184+92 with 0 failures (rc=1 only because tyrell-appTests ran 0 tests). replicantDB 5447787 did NOT compile (NSLock in async context from cloud commit bcc8cc1).
- Fixed inline: replicantDB c5dd78a; swift build rc=0, swift test 1213 tests 0 failures; pushed to backup and fleet.