Model, delegation & spend policy v3.0 — consolidation + fleet deploy
Host: rdmsm4x (lead) · Session: dev-db [982a7e] · Started 13:32 EDT, closed 13:58 EDT, 2026-08-25 America/New_York
Five overlapping CLAUDE.md sections that answered the same question differently were consolidated into one, the three-way lead-tier contradiction was resolved, the unenforceable budget backstop was replaced with a real MCP tool, and the result was deployed to 5 of 6 fleet hosts across all three agent harnesses.
Why
~/.claude/CLAUDE.md carried five sections governing the
same decision, written on five different dates: Spend & scope
discipline (v2.4, 08-13), Model & delegation policy (v2.6, 08-15),
Second opinions via agy (v2.9, 08-22), Unattended escalation to Fable 5
(v1.0, 08-25), External-budget workers (v1.0, 08-25). They
conflicted:
- Lead tier stated three ways. v2.3 "the main thread
never downgrades to save cost" → v2.6 "right-size the lead down to the
lowest capable tier" → Fable v1.0 "default tier is Opus 5 high." A stale
v2.3 artifact was still sitting at
~/.claude/_fleet-model-policy-block.md. - Concurrency caps in three places with different numbers (2 background agents / 2 codex + 1 agy / amended to 1 codex).
- "Unattended" was undefined yet gated the entire Fable 5 pre-authorization — an agent at 03:00 had no way to determine which mode it was in.
- The budget backstop was unenforceable. "When a
usage bucket passes ~75%, stop delegating" — nothing on any host could
read a usage bucket. Searched
~/.claudefor usage/quota/budget telemetry: none exists;/usageis interactive-only. Six Macs were following a rule none could evaluate. - Fleet drift. rdmsm4x was 743 lines (Aug 25); the other four reachable hosts were 623 lines (Aug 23), missing three whole sections. Codex and agy had no model/delegation policy at all — the policy governing cross-vendor delegation had never been given to the vendors doing the work.
Decisions (Rich, this session)
- Lead tier: Opus 5 high is the default, explicitly right-sized down to Sonnet-class for routine work; Max on demonstrated difficulty; Fable 5 trigger-gated only.
- Structure: consolidate all five into one section rather than patching in place.
- Escalation gate: the ask is the
attended/unattended detector. On a trigger the agent always asks; an
answer is binding; a 60s timeout or an unavailable AskUserQuestion
(headless/cron) means unattended → escalate and log. Once unattended is
established, don't re-ask. Rich's added exception: once unattended, the
stuck-after-~3-attempts trigger escalates with no ask at all. Any new
message from Rich resets the session to attended.
- An idle-clock threshold (the 8h variant considered) was rejected: during an active overnight run that clock only grows while the agent works, so a long threshold locks escalation during exactly the hours it exists to cover, and unlocks once Rich is awake.
- Budget: build real usage monitoring into Tyrell today; use a stop-flag file in the interim.
What changed
Tyrell
usage_status MCP tool — built by the
tyrell.app build kickoff session
Coordinated via SendMessage; proposed the contract, that session
implemented and shipped it (99f47e5, 8e49d94).
Contract: stateless re-read per call;
advice: proceed|conserve|hold|unknown
advice_reason;observed_atalways present;unknownfirst-class (used_pctnull, stale degrades to unknown rather than carrying forward);~/.agent-coordination/BUDGET-HOLDforcesholdwith the file's first line as the reason; thresholds (60/75) live in the tool so agents branch on the enum, not a percentage. Schema is bucket-name agnostic so agy/ChatGPT buckets ride the same shape.
Known limit (v1): no live vendor telemetry exists. A
collector writes ~/.tyrell/usage.json; until one exists the
honest answer is unknown = "proceed but announce spend
before fan-outs".
Independently verified here, not taken on report — driving the built binary directly over stdio:
tools/listreturnsusage_statusalongside the four pre-existing tools.- No sources →
{"advice":"unknown","source":"none","observed_at":"2026-08-25T17:51:03Z"}. - With
BUDGET-HOLDpresent →{"advice":"hold","advice_reason":"<file's first line>","source":"manual-flag"}. - Test flag removed afterwards; confirmed absent.
~/.claude/CLAUDE.md →
v3.0
New single section "Model, delegation & spend
(v3.0)" at line 40, 221 lines, seven subsections: tier ladder
(table) · usage_status budget check · 10 delegation rules ·
Fable triggers + gate · worker tiers across vendors · agy specifics ·
the delegation test · revision log v2.3→v3.0.
Removed (content preserved in the consolidation, nothing dropped):
- lines 40–91 Spend & scope + Model & delegation policy
- lines 437–504 the
agy-second-opinions v2.9delimited block (markers included; verified no script injects that block — it was a one-off hand recovery on 08-23, so removing the markers is safe) - lines 653–743 Unattended escalation + External-budget workers
Header bumped v2.8 → v3.0; line-3 preamble updated to name the consolidated policy and to state it binds Codex and agy too.
Fleet deployment
New script:
~/dev/fleet/maintenance/scripts/deploy_model_policy_v3.zsh
(v1.0).
~/.claude/CLAUDE.md: whole-file copy from rdmsm4x, divergence-gated — a host is overwritten only if it has zero lines the pre-deploy canonical lacked; otherwise skipped and reported, never clobbered. All four reachable peers measured zero unique lines.~/.codex/AGENTS.mdand~/.gemini/GEMINI.md: the policy injected as a delimitedfleet-model-delegation-spend v3.0block via the existing_inject_block.py, so peer-written content outside the markers survives byte for byte. The block carries an orientation preamble telling non-Claude harnesses they are a worker tier — §3/§5/§7 bind them, §1/§2/§4 are context.- Per-host backups written before every write:
CLAUDE.md.bak-pre-v3-<ts>,AGENTS.md.bak-pre-v3-<ts>,GEMINI.md.bak-pre-v3-<ts>.
Result — 5/6 hosts, all three context files each:
| Host | CLAUDE.md | AGENTS.md | GEMINI.md |
|---|---|---|---|
| rdmsm4x | canonical (0d28208f48a219b3) | block verified | block verified |
| rdmbair13m5 | updated + sha-verified | block verified | block verified |
| rdmbair15m5 | updated + sha-verified | block verified | block verified |
| rdmpw3265m | updated + sha-verified | block verified | block verified |
| rdmpw3275m | updated + sha-verified | block verified | block verified |
| jdmbair13m5 | PENDING — host unreachable | pending | pending |
Verified independently of the deploy script's own report, by SSH to
rdmbair15m5 and rdmpw3275m: v3.0 header present, exactly
one ## Model, delegation & spend (v3.0 heading,
zero superseded section headings remaining, and
usage_status referenced 5× in each of the three context
files.
Reproduce
# policy present and superseded sections gone, any host
grep -c '^## Model, delegation & spend (v3.0' ~/.claude/CLAUDE.md # -> 1
grep -c '^## Spend & scope discipline\|^## Unattended escalation' ~/.claude/CLAUDE.md # -> 0
# usage_status, both paths (rdmsm4x)
B=~/dev/apps/Tyrell/.build/release/tyrell-mcp
printf '%s\n' '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"v","version":"1"}}}' \
'{"jsonrpc":"2.0","id":3,"method":"tools/call","params":{"name":"usage_status","arguments":{}}}' | $B | tail -1
# re-deploy / catch up a host
zsh ~/dev/fleet/maintenance/scripts/deploy_model_policy_v3.zsh --dry-run
zsh ~/dev/fleet/maintenance/scripts/deploy_model_policy_v3.zsh --host jdmbair13m5Undo
cp ~/.claude/CLAUDE.md.bak-pre-v3-modelpolicy-20260825-135234 ~/.claude/CLAUDE.md # rdmsm4x
# peers: ~/.claude/CLAUDE.md.bak-pre-v3-20260825-135552
# codex/agy: ~/.codex/AGENTS.md.bak-pre-v3-20260825-135552, ~/.gemini/GEMINI.md.bak-pre-v3-20260825-135552Outstanding — owner actions
- jdmbair13m5 has not received v3.0 (asleep at deploy
time; the Tyrell session reported the same host unreachable for its own
deploy). Re-run the deploy with
--host jdmbair13m5when it wakes. Until then that host's agents follow the superseded 08-23 policy. - No usage collector exists, so
usage_statusreturnsunknowneverywhere. The enforceable control today is~/.agent-coordination/BUDGET-HOLD, which is Rich's to set — agents are explicitly forbidden from creating or deleting it. - Second ChatGPT Pro account not authed ([email protected]); codex remains authed to one.
- Grok SuperHeavy joins the worker tiers when live; codex caps revisit at that point.
- No second opinion obtained on this revision — not routed to agy. The changes are policy text plus a reversible, backed-up file deployment, not a data-destroying operation.
No secrets were written to any file, message, or note in this session.