rdmpw3275m changelog — agent-coordinator resumes agy and Claude Code after quota resets
[session resumed 2026-09-21 16:18:27 EDT · written 2026-09-21 16:52 EDT · rdmpw3275m · claude Opus 5 (1M) · session 5739b30e]
Rich, 16:18 EDT: "resume all work when quotas reset for agy/etc".
Where the quotas stood (usage_status, vendor-authoritative, 16:19 EDT)
- Anthropic: 5-hour window 5% (resets 20:20 EDT), 7-day 12% →
proceed. - OpenAI / codex: 7-day 0% (resets 2026-09-26) →
proceed. - Google / agy: no published quota. Provider truth is agy's own 429s.
- xAI / grok: installed but not signed in. That is an authentication step, not a quota.
The only real wall in the fleet was on rdmbair15m5. Seven agy sessions there carried a provider-announced reset of 15:25:05 EDT, which had already passed 56 minutes earlier.
Defect found and fixed: the resumer had been resuming nothing (ISSUE-20260921-09)
com.eastcoastscience.agentcoord (every 300 s on all six
Macs) is the fleet's quota resumer. Its v1.7 held every walled session
as "agy still retrying on its own" while
now - max(last_429, log mtime) < 240 s. agy writes
housekeeping lines to that same log every few minutes, so the mtime
never aged and no walled session was ever resumed after its
reset.
Proof on rdmbair15m5: since the reset, lane 2d005e82 had written 20 log lines (18 http_helpers.go, 1 mcp_manager.go, 1 browser.go) and zero retries, 429s or model calls. agy's own retry loop had ended at attempt 8, at 15:20:19. All 8 active logs showed zero 429s after 15:25, so the quota really had reset.
v1.8 (rdmsm4x:~/dev/fleet
cee2d9c, pushed to fleet and backup): the hold decision
moves into hold_reason(). "Still retrying" now means a 429
within 240 s, or a retry that agy itself scheduled
(retrying in 24.7s) and that is not yet due. mtime is
ignored. DEC-20260918-02 is unchanged.
Result: v1.8 went live on rdmbair15m5 at 16:29:27. Its first pass nudged 2d005e82 (RooReceipt), 25b8267f (CPU monitor) and 29f62094 (ReplicantDB). Within 35 s they had made 5, 7 and 9 model calls, with 0 new 429s.
Extension: Claude Code sessions too (DEC-20260921-01, v1.9 = ad6bc1b)
A Claude Code session that hits its limit stops and waits at its
prompt. It doesn't retry and it doesn't read the bus, so the old bus
broadcast (silent since 2026-09-09) could never unstick one. There were
22 real stops in 10 hub transcripts, all reading
You've hit your session limit · resets 9pm (America/New_York).
- What v1.9 does: it maps
claude --resume <uuid>sessions to their transcripts. A session counts as walled only when its last user or assistant turn is the stop. It ignores the bookkeeping entries Claude appends afterwards, it reads only the session limit as a clock time, and it applies the same guards as agy. - Codex is not covered. No retained rollout holds a real Codex limit stop to build against.
- How it was tested: 28 unit tests, 14 for agy and 14
for Claude. Mutation checks were caught, and restoring the v1.7
condition fails the named regression test. End to end, v1.9's own
TIOCSTI inject of text plus CR was submitted by Claude Code within 1 s.
Session mapping found 7 of 7 live sessions, including the one that
pgrepcannot see. - Dry probe before rollout: 29 interactive Claude sessions on six hosts, 0 at a limit, so the first pass nudges nothing.
- Rollout: v1.9 is installed and verified on
all six hosts. The five spokes ran their first pass at
16:49:17–18 and the hub at 16:49:20 (its pass takes minutes, because it
walks every repo). Every host exited 0 with 0 actions, the installed
binary's sha256 is
bcc502fa6bdd7b20, and the Claude row counts (6, 3, 4, 4, 6, 6) match the probe exactly. The hub's launchd.log has never recorded a traceback.
Also this session
- ISSUE-20260906-02 (apfsd storm), post-reboot check. The disable flag survived the ~15:57 reboot. My morning pass had missed two volumes, and the SSD container had renumbered from disk3 to disk2, so apfsd was back at about one core defragmenting disk2s7, a 48.8 MB SSD volume. I disabled defrag on disk2s7 and disk7s5, and apfsd went back to 0.0%. The sealed system volume was deliberately left alone.
- The hub's coordinator had completed no pass from 15:35:58 until my 16:30 reinstall. A 16:14 dry run was masking that. It completes passes again now.
- Filed ISSUE-20260921-15: stall detection (job 2) has the same mtime root cause and effectively never fires. I did not change it here, because it is alerting behavior nobody asked to change.
- Unexplained side finding:
tmux send-keys ... Enterdid not submit into a Claude Code TUI, while TIOCSTI did. Broadcast to the fleet so anyone relying on a tmux-based resumer checks it. - Tickets: ISSUE-20260921-09 resolved after all six hosts were verified. DEC-20260921-01 resolved. Comment added to codex's TASK-20260906-01. Fleet bus message 20260921-165009-0B5490E0 sent.
How to undo
- Code:
git -C rdmsm4x:~/dev/fleet revert ad6bc1b cee2d9c, then runzsh ~/dev/fleet/agent-coordinator/install_agent_coordinator.zshon each host. - apfs:
sudo diskutil apfs defragment disk2s7 enable(same for disk7s5).
Not done, deliberately
- Codex resume: there is no observed stop format to build against.
- grok: it needs Rich to sign in before it can run at all.
- Idle agy/Claude sessions: DEC-20260918-02 keeps them un-nudged. They are not waiting on a quota.