Model policy v3.1 — task-class tier table, UNATTENDED marker, live usage collector
Host: rdmsm4x · Session: dev-db [982a7e] · 2026-08-25, v3.0 closed 13:58 EDT, v3.1 closed 15:25 EDT America/New_York
Same-day follow-up to v3.0. Rich revised three decisions after seeing the consolidated draft; all three are now live on 5/6 hosts. A usage collector shipped mid-session, so the budget rule went from aspirational to enforced.
What changed from v3.0
§1 — single default replaced by a task-class table
Rich's call: a task-class table rather than one blanket tier, derived from real work rather than asserted. Constraint found and reported: only 12 days of transcripts exist (2026-08-07 → 08-25, 270 sessions on rdmsm4x; earliest anywhere on the fleet is 07-31). A three-month review is not possible — Claude Code does not retain transcripts that long. Transcripts also record which model ran, never which would have sufficed, so the table is calibrated, not proven. Both limits are stated in the policy text itself.
Measured
(~/dev/fleet/maintenance/scripts/tier_analysis.py, archived
and re-runnable):
| Model | Share of assistant turns |
|---|---|
| Sonnet 5 | 41.6% |
| Opus 5 | 34.8% |
| Fable 5 | 22.1% |
| Haiku 4.5 | 1.2% |
| Work class | Sessions | Tool calls | Err/100 | Asst turns/sess | User turns/sess |
|---|---|---|---|---|---|
| ops-shell | 127 | 7,392 | 3.6 | 121 | 3.5 |
| implementation | 48 | 6,606 | 3.0 | 274 | 13.4 |
| mixed | 83 | 4,966 | 3.8 | 106 | 3.1 |
| investigation | 10 | 395 | 2.0 | 66 | 1.1 |
Decisive finding: tool-error rate is flat (2.0–3.8 per 100) across every class, so difficulty does not separate them. What separates them is reversibility and how little supervision they run under. The table was therefore keyed on blast radius and autonomy, and the policy explicitly forbids re-keying it on difficulty. Tool mix is 58% Bash, 13% Read, 9% Edit.
Adopted table: investigation → Haiku · ops-shell → Sonnet 5, upgrading on the 2nd failed cycle (Rich's amendment; it is the highest-volume least-supervised class) · implementation → Opus 5 high · irreversible (signing, schema, destructive, release) → Opus 5 max · unattended+stuck → Fable 5.
Noted in policy: Haiku at 1.2% against a 58%-Bash workload is the largest unrealised saving.
§4 — gate changed to an explicit marker
Replaces the ask-with-timeout detector agreed earlier in the same
session. Now: test -f ~/.agent-coordination/UNATTENDED.
Present ⇒ escalate without asking, log the trigger. Absent ⇒ attended ⇒
ask, and a timed-out or unavailable ask is never an implicit
yes — stay at the current tier. Agents never create or delete
the marker. The accepted tradeoff (a forgotten touch = no
escalation that night) is written into the policy rather than left
implicit.
§2 — collector
shipped, so unknown is now a fault
The tyrell.app build kickoff session shipped
scripts/usage_collector.zsh, aggregating per-message token
usage from ~/.claude/projects/**/*.jsonl into rolling
session_5h / weekly_7d buckets.
Verified independently here, driving the binary over stdio —
and it had already moved past what the peer reported
(conserve):
advice: hold — "worst bucket 99% used — budget backstop (finish inline, no delegation)"
session_5h : 14,786,640 / 15,000,000 tokens = 98.6% (fable 8.28M · sonnet 4.17M · opus 2.34M)
weekly_7d : 128,980,626 / 200,000,000 tokens = 64.5% (opus 67.3M · fable 54.8M · sonnet 6.2M)
observed_at 2026-08-25T19:21:40Z · stale_after 900s · source local-transcripts
The remainder of the session honored that hold:
no delegation, no subagents, all work inline.
Critical caveat written into §2 rather than glossed: the ceilings in
~/.tyrell/usage-limits.json are provisional, set
from observed volume — not vendor-published plan limits. No
programmatic vendor quota read exists anywhere. A percentage is only as
right as its ceiling. advice is directional; act on it, and
say the ceiling is provisional.
Naming decision made rather than asked: kept
BUDGET-HOLD, rejected the proposed
BUDGET_BRAKE — the MCP tool already reads that exact path
and a second name recreates the two-competing-paths problem the tool
exists to prevent. Semantics unchanged.
New datapoint Rich did not have when he chose to leave Fable gating alone: Fable is ~42% of weekly tokens versus 22.1% of turns. Logged for the scheduled re-measure, not re-litigated.
Deploy defect found and fixed
The first v3.1 deploy correctly refused to write all four
peers, reporting 215 diverged lines each. That was the
divergence gate working, on a stale baseline: it compared peers against
the pre-v3 backup, which legitimately lacks the v3.0 section
deployed 90 minutes earlier. Fixed by tracking
~/.claude/.CLAUDE.md.last-deployed, written after each
successful run and seeded from a peer's post-v3.0 file. Re-ran clean.
No host was overwritten while the check was wrong — the
gate failed safe, which is what it was built to do.
Result — 5/6 hosts, sha 1306d4dd6406e707
| Host | CLAUDE.md | AGENTS.md | GEMINI.md |
|---|---|---|---|
| rdmsm4x | canonical | block verified | block verified |
| rdmbair13m5 | v3.1 verified | block verified | block verified |
| rdmbair15m5 | v3.1 verified | block verified | block verified |
| rdmpw3265m | v3.1 verified | block verified | block verified |
| rdmpw3275m | v3.1 verified | block verified | block verified |
| jdmbair13m5 | PENDING — unreachable | pending | pending |
Independently verified on rdmbair13m5 and rdmpw3265m by SSH: v3.1
header, blast-radius table present, UNATTENDED marker
present, collector-fault language present, and the delimited block live
in both codex and agy context files.
Reproduce / undo
python3 ~/dev/fleet/maintenance/scripts/tier_analysis.py # regenerate the tier evidence
zsh ~/dev/fleet/maintenance/scripts/deploy_model_policy_v3.zsh --dry-run
zsh ~/dev/fleet/maintenance/scripts/deploy_model_policy_v3.zsh --host jdmbair13m5 # catch up
# undo on any peer: ~/.claude/CLAUDE.md.bak-pre-v3-20260825-152* (per-host, written before each write)Outstanding
- jdmbair13m5 still has no v3.0 or v3.1 — unreachable all session. Its agents follow the 08-23 policy.
- Provisional ceilings in
~/.tyrell/usage-limits.json— Rich to tune; percentages inherit their error. - Re-measure ~2026-09-01 to see whether Fable's share falls under §4 gating.
- No independent review of v3.1 (not routed to agy). Policy text plus a reversible, backed-up, divergence-gated deployment.
No secrets were written to any file, message, or note.