Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260909-1300-control-plane-monitoring-only-stall-recovery

rdmsm4x-changelog-20260909-1300-control-plane-monitoring-only-stall-recovery

Fleet execution stall (since ~2026-09-04) unblocked: control-plane automation set to monitoring only, cloud routines off, hub synced with GitHub mirrors. Written 2026-09-09 13:00:12 EDT by claude@rdmsm4x (Fable 5.1) session_01L5FqUPjVWGsMbP4UhAqAa6.

DEC-20260909-01 — Fleet control plane → MONITORING ONLY; work local, no cloud, no containers (2026-09-09 13:00:12 EDT · rdmsm4x)

Rich, live, 2026-09-09 12:4x EDT: "we are seemingly stuck on execution… if stuck on Tyrell, temporarily unhook it and put it back to monitoring mode vs. control plane… un-do containers… we should be working locally… things have seemed stalled since right before Labor Day." Lead: claude@rdmsm4x session_01L5FqUPjVWGsMbP4UhAqAa6 (Fable 5.1, oversight tier).

Root causes found (all installed 2026-09-04/05, i.e. right before Labor Day):

  1. fleet-load-guard (90 s) broadcast a fleet-wide SLOWDOWN directive every ~30–60 min and a recovery notice in between — flip-flop all morning (12:10, 12:35 today; 10:11, 11:28 earlier). Every agent that read the bus was told to stagger dispatches 45–60 s and cut parallelism, repeatedly.
  2. usage_status advice = conserve (Anthropic weekly 72 %, resets 2026-09-11 10:00 EDT) — by CLAUDE.md that forbids every fan-out on every agent. OpenAI is at 0 % with 2 500 credits and was unused.
  3. Cloud sessions/routines: 5 idle cloud sessions (Tyrell, ReplicantDB, ecs0lib, RTTy ×2) and 20 routines (2 enabled) were doing "oversight" from Ubuntu containers with no Swift toolchain, waiting on a disconnected hub bridge, and merging PRs to the GitHub mirrors uncompiled ("MERGE_AND_INTEGRATION rule 3"). Hub checkouts had no origin; the merges never reached the hub or the fleet bare store.
  4. claude-quota-monitor ran ssh host "claude agents --json" with a Python timeout that kills only the local ssh — 121 orphaned probes on rdmbair13m5 (one every 6 min).
  5. stormguard / fleet-realtime-guard kill processes; resume-dispatcher (RunAtLoad) respawned idle sessions into Terminal on every boot. Boot storm at 12:16 today: load1 143 on 16 cores.

Actions taken (all reversible):

Interpretation of conserve during this recovery: it means no Claude fan-out beyond one worker at a time and route bounded implementation to codex exec; it does NOT mean idle. Rich's explicit unblock instruction today outranks the gate. Re-evaluate at the weekly reset 2026-09-11 10:00 EDT.

Left alone / needs Rich: apps/macOS-ramDisk-app_eval is a second checkout of the same GitHub repo as apps/ramdisk, diverged with 5 overlapping files (INC-20260905-01 recovery commit) — not merged. Dirty trees not touched: fleet (570), sites/dev.ecs0.net (102), ai/LLM (106), docs (48), issues (15). UNATTENDED marker still present (Rich's).

To re-enable the control plane later: launchctl enable gui/$(id -u)/<label> && launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/<label>.plist per job; fix the quota monitor's remote timeout (kill the remote side, e.g. ssh host 'timeout 5 claude agents --json') before re-enabling it; re-enable routines at https://claude.ai/code/routines.

Undo

Follow-through (2026-09-09 13:15:23 EDT)