rdmsm4x-changelog-20260903-1526-spoke-preflight-gate-and-publish-job
rdmsm4x-changelog-20260903-1526-spoke-preflight-gate-and-publish-job
[2026-09-03 15:26:16 EDT · rdmsm4x · claude@rdmsm4x session 158c9525 (notes-republish-pending-changelog)]
Four fleet-maintenance defects reported by codex sessions on jdmbair13m5 and rdmbair13m5 at 14:52-14:55 EDT are fixed at the source, distributed to all six Macs, and verified by re-running the exact reported commands on the spokes; the dev.dataroo.net publish job is now bounded and timed per stage, and the agent board it feeds regenerates again. Ticket: TASK-20260903-30 (resolved).
Scope
- Hosts: all six (rdmsm4x, rdmbair13m5, rdmbair15m5, jdmbair13m5, rdmpw3265m, rdmpw3275m).
- Repos: fleet/maintenance submodule (e74d218), lib/agentkit (5eb5bfa, 9417907), both pushed to GitHub backup and git.ecs0.net fleet.
- Files outside git:
~/.agent-coordination/agent_msg.zsh(six hosts),~/bin/ticket(five spokes),~/dev/lib/agentkit/lib/plain copies (four spokes),~/scripts/agent_checkpoint.zsh(six hosts) plus the six untracked snapshots under~/dev/fleet/tooling/host-scripts/,~/.agent-coordination/AGENT_COORDINATION.md,~/.claude/skills/fleet-ticketing/SKILL.md,~/dev/scripts/publish_site.zsh(untracked in the macos-scripts repo).
What changed and why
- Bus inbox 59 s -> <1 s.
cmd_inboxran one awk per front-matter field per message (about seven processes per message) over 3,396 messages. v1.4.0 reads all front matter in one awk pass and loads the seen-list once. Output proven byte-identical on the live mailbox (2,850 rows, diff empty;--all3,395,--unread1,451). Installed with checksum guards; the previous copy is kept beside it as.bak-v1.3.0-20260903. Spoke timings after install: 0-2 s. - Ticket CLI on pointer-only spokes. The retirement
removed
~/dev/issuesfrom spokes, so the mandatory preflight returned BROKEN on every spoke and the fleet-ticketing skill's "browsing stays local" was no longer true.~/bin/ticketon each spoke is now a 30-line proxy that runs every subcommand on rdmsm4x with the spoke's identity, using the same env contract as the CLI's ownproxy_to_authority. Verified per spoke:claims,show, and the conformance probe. rdmbair15m5 had a dangling symlink at that path (moved to~/bin/ticket.dangling-symlink-20260903, not deleted). - agentkit preflight v1.1.1 accepts a ticket proxy
found on PATH before declaring the CLI missing. v1.1 (5eb5bfa) was
wrong: a perl substitution with
|as delimiter turned the escaped||into an empty alternation and inserted the block at file top; all three suites still passed because both paths resolve on rdmsm4x. Caught only by re-running the codex repro on jdmbair13m5; repaired in 9417907 with exact-string edits.lib/copied without.gitto the four spokes that lacked it; the rdmbair13m5 clone fast-forwarded via its origin. - agent_checkpoint.zsh used Bash-only
${VAR@Q}under a zsh shebang (exit 1 "bad substitution" at the JSON heredoc; and even under Bash it would have emitted single-quoted, non-JSON strings). Replaced with a zshjson_strescaper and(qq)quoting for the ssh branch. Validated with hostile input through python's json parser; the rule-30 fallback on jdmbair13m5 now writes valid JSON with the right host and agent. - Policy text. AGENT_COORDINATION.md section 7 gained
a pointer-only-spoke paragraph (backup
AGENT_COORDINATION.md.bak-pre-spoke-preflight-20260903-151956); the ticketing skill notes the proxy. - publish_site.zsh v1.1. Every stage now runs under
coreutils
timeoutand its rc/seconds are appended to~/dev/data/publish.logas astagesline;gen_agent_board.pywas belowexitand had not run since 2026-08-29 (the board Rich reads was five days stale). First v1.1 run failed rc 127 at 15:21 because zsh does not word-split${TMO:+$TMO 900}; fixed to an array wrapper. Verified run: rc 0 in 69 s (merge 0 s, projects 2 s, site 6 s, ecs0 61 s, board 0 s);~/dataroo.net/wiki/agents.htmlregenerated 15:24. Backup:publish_site.zsh.bak-20260903-152118.
Measured after the fix (login shell, 60 s cap, as codex ran it)
| Host | Preflight | Time |
|---|---|---|
| jdmbair13m5 | rc 1 (ACT: aging HIGH messages) | 2 s |
| rdmbair13m5 | rc 1 | 1 s |
| rdmbair15m5 | rc 1 | 2 s |
| rdmsm4x | rc 1 | ~3 s |
rc 1 is the intended signal (unacked HIGH messages aging), not BROKEN. The fleet-wide non-acking of HIGH messages (685 for claude@rdmsm4x, oldest 11.8 days) is a separate, pre-existing finding (nightly bus-unread-high).
Undo
- Bus:
cp ~/.agent-coordination/agent_msg.zsh.bak-v1.3.0-20260903 ~/.agent-coordination/agent_msg.zshon any host;git -C ~/dev/fleet/maintenance revert e74d218. - agentkit:
git -C ~/dev/lib/agentkit revert 9417907 5eb5bfa; re-copylib/to spokes. - Proxy: remove
~/bin/ticketon a spoke (on rdmbair15m5 restore the moved symlink only if~/dev/issuesreturns). - Checkpoint:
~/dev/fleet/tooling/host-scripts/*/agent_checkpoint.zshwere updated in place; the original is in this session's scratchpad and in the spoke copies' history (sha 17be5e91). - publish_site.zsh:
cp ~/dev/scripts/publish_site.zsh.bak-20260903-152118 ~/dev/scripts/publish_site.zsh.
Left open
- Root cause of the 2-6 h gaps in devsite delivery is unproven; the
new
stageslog line attributes the next one. Not sleep (pmset shows none); not stormguard (it does not stop LaunchAgents). - Fleet-wide HIGH-message acking discipline (nightly bus-unread-high, 612 unread).
~/scriptshas no distribution mechanism; the sixfleet/tooling/host-scriptssnapshots are untracked.
Secrets
None written or referenced.