Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260825-1525-model-policy-v3.1-task-class-tiers-unattended-marker-usage-collector

Model policy v3.1 — task-class tier table, UNATTENDED marker, live usage collector

Host: rdmsm4x · Session: dev-db [982a7e] · 2026-08-25, v3.0 closed 13:58 EDT, v3.1 closed 15:25 EDT America/New_York

Same-day follow-up to v3.0. Rich revised three decisions after seeing the consolidated draft; all three are now live on 5/6 hosts. A usage collector shipped mid-session, so the budget rule went from aspirational to enforced.

What changed from v3.0

§1 — single default replaced by a task-class table

Rich's call: a task-class table rather than one blanket tier, derived from real work rather than asserted. Constraint found and reported: only 12 days of transcripts exist (2026-08-07 → 08-25, 270 sessions on rdmsm4x; earliest anywhere on the fleet is 07-31). A three-month review is not possible — Claude Code does not retain transcripts that long. Transcripts also record which model ran, never which would have sufficed, so the table is calibrated, not proven. Both limits are stated in the policy text itself.

Measured (~/dev/fleet/maintenance/scripts/tier_analysis.py, archived and re-runnable):

Model Share of assistant turns
Sonnet 5 41.6%
Opus 5 34.8%
Fable 5 22.1%
Haiku 4.5 1.2%
Work class Sessions Tool calls Err/100 Asst turns/sess User turns/sess
ops-shell 127 7,392 3.6 121 3.5
implementation 48 6,606 3.0 274 13.4
mixed 83 4,966 3.8 106 3.1
investigation 10 395 2.0 66 1.1

Decisive finding: tool-error rate is flat (2.0–3.8 per 100) across every class, so difficulty does not separate them. What separates them is reversibility and how little supervision they run under. The table was therefore keyed on blast radius and autonomy, and the policy explicitly forbids re-keying it on difficulty. Tool mix is 58% Bash, 13% Read, 9% Edit.

Adopted table: investigation → Haiku · ops-shell → Sonnet 5, upgrading on the 2nd failed cycle (Rich's amendment; it is the highest-volume least-supervised class) · implementation → Opus 5 high · irreversible (signing, schema, destructive, release) → Opus 5 max · unattended+stuck → Fable 5.

Noted in policy: Haiku at 1.2% against a 58%-Bash workload is the largest unrealised saving.

§4 — gate changed to an explicit marker

Replaces the ask-with-timeout detector agreed earlier in the same session. Now: test -f ~/.agent-coordination/UNATTENDED. Present ⇒ escalate without asking, log the trigger. Absent ⇒ attended ⇒ ask, and a timed-out or unavailable ask is never an implicit yes — stay at the current tier. Agents never create or delete the marker. The accepted tradeoff (a forgotten touch = no escalation that night) is written into the policy rather than left implicit.

§2 — collector shipped, so unknown is now a fault

The tyrell.app build kickoff session shipped scripts/usage_collector.zsh, aggregating per-message token usage from ~/.claude/projects/**/*.jsonl into rolling session_5h / weekly_7d buckets.

Verified independently here, driving the binary over stdio — and it had already moved past what the peer reported (conserve):

advice: hold — "worst bucket 99% used — budget backstop (finish inline, no delegation)"
session_5h : 14,786,640 / 15,000,000 tokens = 98.6%   (fable 8.28M · sonnet 4.17M · opus 2.34M)
weekly_7d  : 128,980,626 / 200,000,000 tokens = 64.5% (opus 67.3M · fable 54.8M · sonnet 6.2M)
observed_at 2026-08-25T19:21:40Z · stale_after 900s · source local-transcripts

The remainder of the session honored that hold: no delegation, no subagents, all work inline.

Critical caveat written into §2 rather than glossed: the ceilings in ~/.tyrell/usage-limits.json are provisional, set from observed volume — not vendor-published plan limits. No programmatic vendor quota read exists anywhere. A percentage is only as right as its ceiling. advice is directional; act on it, and say the ceiling is provisional.

Naming decision made rather than asked: kept BUDGET-HOLD, rejected the proposed BUDGET_BRAKE — the MCP tool already reads that exact path and a second name recreates the two-competing-paths problem the tool exists to prevent. Semantics unchanged.

New datapoint Rich did not have when he chose to leave Fable gating alone: Fable is ~42% of weekly tokens versus 22.1% of turns. Logged for the scheduled re-measure, not re-litigated.

Deploy defect found and fixed

The first v3.1 deploy correctly refused to write all four peers, reporting 215 diverged lines each. That was the divergence gate working, on a stale baseline: it compared peers against the pre-v3 backup, which legitimately lacks the v3.0 section deployed 90 minutes earlier. Fixed by tracking ~/.claude/.CLAUDE.md.last-deployed, written after each successful run and seeded from a peer's post-v3.0 file. Re-ran clean. No host was overwritten while the check was wrong — the gate failed safe, which is what it was built to do.

Result — 5/6 hosts, sha 1306d4dd6406e707

Host CLAUDE.md AGENTS.md GEMINI.md
rdmsm4x canonical block verified block verified
rdmbair13m5 v3.1 verified block verified block verified
rdmbair15m5 v3.1 verified block verified block verified
rdmpw3265m v3.1 verified block verified block verified
rdmpw3275m v3.1 verified block verified block verified
jdmbair13m5 PENDING — unreachable pending pending

Independently verified on rdmbair13m5 and rdmpw3265m by SSH: v3.1 header, blast-radius table present, UNATTENDED marker present, collector-fault language present, and the delimited block live in both codex and agy context files.

Reproduce / undo

python3 ~/dev/fleet/maintenance/scripts/tier_analysis.py          # regenerate the tier evidence
zsh ~/dev/fleet/maintenance/scripts/deploy_model_policy_v3.zsh --dry-run
zsh ~/dev/fleet/maintenance/scripts/deploy_model_policy_v3.zsh --host jdmbair13m5   # catch up
# undo on any peer: ~/.claude/CLAUDE.md.bak-pre-v3-20260825-152* (per-host, written before each write)

Outstanding

  1. jdmbair13m5 still has no v3.0 or v3.1 — unreachable all session. Its agents follow the 08-23 policy.
  2. Provisional ceilings in ~/.tyrell/usage-limits.json — Rich to tune; percentages inherit their error.
  3. Re-measure ~2026-09-01 to see whether Fable's share falls under §4 gating.
  4. No independent review of v3.1 (not routed to agy). Policy text plus a reversible, backed-up, divergence-gated deployment.

No secrets were written to any file, message, or note.