Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260825-1905-claude-degradation-rootcause-and-fleet-remediation

rdmsm4x - remediation - claude-degradation-rootcause-and-fleet-remediation - claude - claude-issues-082526 - fleet - 20260825-1905

Session: 2026-08-25 18:40:56 EDT to 19:05 EDT Ran on: rdmsm4x · Changed state on: rdmsm4x, rdmbair13m5, rdmbair15m5, rdmpw3265m, rdmpw3275m (jdmbair13m5 pending, self-heals on boot)

Claude sessions had become slow, stopped following instructions, conflated unrelated threads, and burned roughly an extra $100. Root cause was three local configuration faults, not the model and not permissions.defaultMode. All three fixed and verified fleet-wide.

Scope

Five of six fleet hosts changed and verified. jdmbair13m5 was offline and converges automatically via the configsync launchd job (1800s interval) once booted — CLAUDE.md, the new skill, settings keys, and sysmon.zsh are all now reconciled, so no manual replay is needed.

Deliberately left alone: permissions.defaultMode (verified bypassPermissions on all five reachable hosts before and after; the fleet patcher aborts without writing if it would move). Per-host settings keys — theme, editorMode, statusLine, hooks, skillOverrides, worktree — were not synced. The frontend-design and hookify plugins remain enabled. ~/.agent-coordination/BUDGET-HOLD and UNATTENDED were not created or removed; those are Rich's controls.

Root causes found

  1. Runaway monitor. ~/dev/apps/Tyrell/scripts/sysmon.zsh fired claude --model sonnet -p on any process whose RSS merely rose for six consecutive samples. In 24h on rdmsm4x: 48 evidence captures, 44 model calls, 0 REAL verdicts, ~130K context each (~5.7–8M tok/day). It also watched claude itself, so triage sessions filed leak reports about triage sessions. Peers were NOT affected (0 captures each) — the loop ran only on rdmsm4x.
  2. 1M-token default context. model: "claude-fable-5[1m]" on rdmsm4x, opus[1m] on peers. Overnight sessions reached 997K/2121 turns, 774K/972 turns, 728K/949 turns — where instruction-following degrades and unrelated threads bleed together.
  3. Instruction-stack bloat. Session-start context tripled in 11 days: ~54K (08-14) → ~69K (08-17) → ~75K (08-24) → 130–180K (08-25). The 59KB CLAUDE.md alone was 14,760 tok/session. The superpowers SessionStart hook injected a mandatory "YOU DO NOT HAVE A CHOICE" skill directive into every session — the direct cause of the halting behaviour.

Not a cause, ruled out: permissions.defaultMode (correct everywhere) and the Claude Code install (2.1.243, executable, on the claude-code@latest fast cask, 2.1.245 published). /doctor would have returned clean.

Genuinely upstream: Anthropic status page confirms elevated error rates for Opus 5 and Fable 5, 2026-08-23 21:50 PT → 2026-08-24 00:36 PT (00:50–03:36 EDT), inside the overnight run.

Files created

Path What
~/.claude/skills/fleet-operations/SKILL.md New on-demand skill holding the standing fleet reference moved out of CLAUDE.md
~/.claude/skills/fleet-operations/references/*.md 7 reference files + the full prior CLAUDE.md, verbatim
~/dataroo.net/wiki/claude-issues-082526.html Published diagnosis + remediation record, self-contained HTML
~/.agent-coordination/canonical/patch_settings.py Canonical settings key-patcher used by the reconciler

Files modified

Commands run

# diagnosis
python3 <ctx analysis over ~/.claude/projects/**/*.jsonl>   # session-start + peak context
mcp tyrell usage_status                                     # advice=hold, weekly 86%
curl https://status.claude.com/                             # upstream incident window

# remediation
python3 <patch settings.json: model + plugins, asserts defaultMode unchanged>
zsh -n ~/dev/apps/Tyrell/scripts/sysmon.zsh                 # syntax gate before deploy
python3 ~/dev/fleet/maintenance/scripts/_divergence_check.py \
        <old-canonical> <new-canonical> <lastdep> <baks>    # rehearsal: returned "OK 0"
zsh ~/dev/fleet/maintenance/scripts/fleet_config_reconcile.zsh --dry-run
zsh ~/dev/fleet/maintenance/scripts/fleet_config_reconcile.zsh
zsh ~/dev/scripts/publish_site.zsh

Verification performed

The trap worth recording

Trimming CLAUDE.md made canonical dramatically shorter, so every peer suddenly held ~46KB of lines canonical lacked — precisely the shape the reconciler's divergence gate blocks, which would have silently stranded the whole fleet. It was safe only because .CLAUDE.md.last-deployed still held the old 59KB version, putting those lines in the known-history set. This was rehearsed before writing with a direct _divergence_check.py run (OK 0) and a dry run showing diverged=0. A shortening change to a divergence-gated file must always be rehearsed that way.

Outstanding owner actions

  1. Boot jdmbair13m5. It converges on its own within 30 minutes of coming online; confirm with cat ~/.agent-coordination/config-sync-state.json (it should leave offline_pending).
  2. Request credits. Cannot be done by an agent. Evidence: an automated loop consuming ~5.7–8M tok/day at a 0% actionable rate (60 sessions on 08-25), plus the confirmed Opus 5 / Fable 5 incident window above. Both independently verifiable from ~/.claude/projects/**/*.jsonl and ~/Library/Logs/Tyrell/sysmon/state/benign.log.
  3. Optional: brew upgrade --cask claude-code@latest (2.1.243 → 2.1.245), and re-assert the exec bit afterwards per the known cask-upgrade trap.
  4. Working pattern: one project per session, /clear between projects. A 2,000-turn session is a bug, not a long day.

Rollback

Every change is reversible from the backups listed above. Re-enable a plugin by flipping its boolean in enabledPlugins. Run a deliberate AI-triage shift with SYSMON_AI_TRIAGE=1 zsh ~/dev/apps/Tyrell/scripts/sysmon.zsh.

Published record: https://dev.dataroo.net/claude-issues-082526.html (Basic auth, by design).