rdmsm4x - remediation - claude-degradation-rootcause-and-fleet-remediation - claude - claude-issues-082526 - fleet - 20260825-1905
Session: 2026-08-25 18:40:56 EDT to 19:05 EDT
Ran on: rdmsm4x · Changed state
on: rdmsm4x, rdmbair13m5,
rdmbair15m5, rdmpw3265m,
rdmpw3275m (jdmbair13m5 pending, self-heals on boot)
Claude sessions had become slow, stopped following instructions,
conflated unrelated threads, and burned roughly an extra $100. Root
cause was three local configuration faults, not the model and not
permissions.defaultMode. All three fixed and verified
fleet-wide.
Scope
Five of six fleet hosts changed and verified. jdmbair13m5 was
offline and converges automatically via the
configsync launchd job (1800s interval) once booted —
CLAUDE.md, the new skill, settings keys, and sysmon.zsh are all now
reconciled, so no manual replay is needed.
Deliberately left alone:
permissions.defaultMode (verified
bypassPermissions on all five reachable hosts before and
after; the fleet patcher aborts without writing if it would move).
Per-host settings keys — theme, editorMode,
statusLine, hooks,
skillOverrides, worktree — were not synced.
The frontend-design and hookify plugins remain
enabled. ~/.agent-coordination/BUDGET-HOLD and
UNATTENDED were not created or removed; those are Rich's
controls.
Root causes found
- Runaway monitor.
~/dev/apps/Tyrell/scripts/sysmon.zshfiredclaude --model sonnet -pon any process whose RSS merely rose for six consecutive samples. In 24h on rdmsm4x: 48 evidence captures, 44 model calls, 0 REAL verdicts, ~130K context each (~5.7–8M tok/day). It also watchedclaudeitself, so triage sessions filed leak reports about triage sessions. Peers were NOT affected (0 captures each) — the loop ran only on rdmsm4x. - 1M-token default context.
model: "claude-fable-5[1m]"on rdmsm4x,opus[1m]on peers. Overnight sessions reached 997K/2121 turns, 774K/972 turns, 728K/949 turns — where instruction-following degrades and unrelated threads bleed together. - Instruction-stack bloat. Session-start context tripled in 11 days: ~54K (08-14) → ~69K (08-17) → ~75K (08-24) → 130–180K (08-25). The 59KB CLAUDE.md alone was 14,760 tok/session. The superpowers SessionStart hook injected a mandatory "YOU DO NOT HAVE A CHOICE" skill directive into every session — the direct cause of the halting behaviour.
Not a cause, ruled out:
permissions.defaultMode (correct everywhere) and the Claude
Code install (2.1.243, executable, on the
claude-code@latest fast cask, 2.1.245 published).
/doctor would have returned clean.
Genuinely upstream: Anthropic status page confirms elevated error rates for Opus 5 and Fable 5, 2026-08-23 21:50 PT → 2026-08-24 00:36 PT (00:50–03:36 EDT), inside the overnight run.
Files created
| Path | What |
|---|---|
~/.claude/skills/fleet-operations/SKILL.md |
New on-demand skill holding the standing fleet reference moved out of CLAUDE.md |
~/.claude/skills/fleet-operations/references/*.md |
7 reference files + the full prior CLAUDE.md, verbatim |
~/dataroo.net/wiki/claude-issues-082526.html |
Published diagnosis + remediation record, self-contained HTML |
~/.agent-coordination/canonical/patch_settings.py |
Canonical settings key-patcher used by the reconciler |
Files modified
~/dev/apps/Tyrell/scripts/sysmon.zsh— v1.1 → v1.2. AddedLEAK_MIN_DELTA_KB=25600(was: any RSS increase counted); extendedEXEMPTwith the 100%-false-positive processes and withclaude/Claude/tyrelldto break the self-triage loop; made Tier-2 model triage opt-in behindSYSMON_AI_TRIAGE=1(evidence capture still always runs, it is free). Backups:~/.claude/backups/sysmon.zsh.bak-20260825-1850(rdmsm4x),~/dev/apps/Tyrell/scripts/sysmon.zsh.bak-v11-20260825(peers).~/.claude/settings.json(all 5 hosts) —model→claude-opus-5(dropping[1m]);modelSettings["claude-opus-5"].effortLevel→high; disabledsuperpowers,ralph-loop,learning-output-style,explanatory-output-style. Backups:~/.claude/settings.json.bak-modelfix-<TS>per host,~/.claude/backups/settings.json.bak-20260825-1850on rdmsm4x.~/.claude/CLAUDE.md— v3.1 → v4.0. 59,042 → 12,596 bytes (14,760 → 3,149 tok). Repackaging, not a rewrite: all 22 sections verified accounted for, zero unaccounted. Backups:~/.claude/skills/fleet-operations/references/CLAUDE-md-full-20260825.md,~/dev/LLM/Claude/changelogs/rdmsm4x-CLAUDE.md-pre-trim-20260825.md,~/.claude/CLAUDE.md.bak-pre-v4-20260825-1850.~/dev/fleet/maintenance/scripts/fleet_config_reconcile.zsh— v1.0 → v1.1. Now also syncs thefleet-operationsskill, patches settings.json keys (never the file), and convergessysmon.zsh. Backup:~/.claude/backups/fleet_config_reconcile.zsh.bak-20260825-1900.
Commands run
# diagnosis
python3 <ctx analysis over ~/.claude/projects/**/*.jsonl> # session-start + peak context
mcp tyrell usage_status # advice=hold, weekly 86%
curl https://status.claude.com/ # upstream incident window
# remediation
python3 <patch settings.json: model + plugins, asserts defaultMode unchanged>
zsh -n ~/dev/apps/Tyrell/scripts/sysmon.zsh # syntax gate before deploy
python3 ~/dev/fleet/maintenance/scripts/_divergence_check.py \
<old-canonical> <new-canonical> <lastdep> <baks> # rehearsal: returned "OK 0"
zsh ~/dev/fleet/maintenance/scripts/fleet_config_reconcile.zsh --dry-run
zsh ~/dev/fleet/maintenance/scripts/fleet_config_reconcile.zsh
zsh ~/dev/scripts/publish_site.zshVerification performed
- CLAUDE.md identical fleet-wide —
shasum -a256compared independently per host after the reconciler ran:b436f1dcf61734a9on all five. defaultModeintact — read back asbypassPermissionson all five after the patch.- sysmon parses and exempts correctly —
zsh -npassed on every host; exemption behaviour table-tested against the actual list of processes flagged today. - No section lost from CLAUDE.md — programmatic check
of all 22 old
##headers against the new file plus the reference set: 0 unaccounted. - Published page serves — HTTP 200, 22,149 bytes;
cross-host verified by fetching from rdmbair15m5 over
Tailscale (not a localhost-only claim); zero external dependencies
confirmed; auto-indexed into
directory.htmlafter a forced publish cycle. - Reconciler idempotent — second full run reported
current=5 updated=0 diverged=0.
The trap worth recording
Trimming CLAUDE.md made canonical dramatically
shorter, so every peer suddenly held ~46KB of lines
canonical lacked — precisely the shape the reconciler's divergence gate
blocks, which would have silently stranded the whole fleet. It was safe
only because .CLAUDE.md.last-deployed still held the old
59KB version, putting those lines in the known-history set. This was
rehearsed before writing with a direct
_divergence_check.py run (OK 0) and a dry run
showing diverged=0. A shortening change to a
divergence-gated file must always be rehearsed that way.
Outstanding owner actions
- Boot jdmbair13m5. It converges on its own within 30
minutes of coming online; confirm with
cat ~/.agent-coordination/config-sync-state.json(it should leaveoffline_pending). - Request credits. Cannot be done by an agent.
Evidence: an automated loop consuming ~5.7–8M tok/day at a 0% actionable
rate (60 sessions on 08-25), plus the confirmed Opus 5 / Fable 5
incident window above. Both independently verifiable from
~/.claude/projects/**/*.jsonland~/Library/Logs/Tyrell/sysmon/state/benign.log. - Optional:
brew upgrade --cask claude-code@latest(2.1.243 → 2.1.245), and re-assert the exec bit afterwards per the known cask-upgrade trap. - Working pattern: one project per session,
/clearbetween projects. A 2,000-turn session is a bug, not a long day.
Rollback
Every change is reversible from the backups listed above. Re-enable a
plugin by flipping its boolean in enabledPlugins. Run a
deliberate AI-triage shift with
SYSMON_AI_TRIAGE=1 zsh ~/dev/apps/Tyrell/scripts/sysmon.zsh.
Published record:
https://dev.dataroo.net/claude-issues-082526.html (Basic
auth, by design).