rdmsm4x-changelog-20260903-1537-bus-v1.5.0-addressing-identity-take-roster
Fleet bus addressing fixed: claude@<host> /
@<host> now route to a group of
session identities, one of which can take a request
while the others stand down — deployed to all six Macs, sha-verified,
with the ticket-ledger half in a PR awaiting merge.
Session: claude@rdmsm4x/55b61233 (dev-33) · 2026-09-03 15:08–15:37 EDT · ticket FEAT-20260903-08 · directed by Rich 15:08 EDT.
Scope
All six hosts (rdmsm4x hub; rdmbair13m5,
rdmbair15m5, jdmbair13m5,
rdmpw3265m, rdmpw3275m). No app source; no
permissions.defaultMode; no credentials.
Files touched
| Host(s) | Path | Change | Backup |
|---|---|---|---|
| 6/6 | ~/.agent-coordination/agent_msg.zsh |
v1.4.0 → v1.5.0 (sha
5c2fdde422432e4c…) |
agent_msg.zsh.bak-v1.4.0-<ts>
beside it on every host |
| 6/6 | ~/.agent-coordination/git-hooks/prepare-commit-msg |
reads CLAUDE_CODE_SESSION_ID
(sha 1ec70f1964bbc9d2…) |
prepare-commit-msg.bak-<ts>
beside it |
| hub | ~/.agent-coordination/AGENT_COORDINATION.md |
appended §9 Addresses are groups; identities are sessions; offers are accepted, not assigned | .bak-pre-section9-<ts>;
reconciler propagates |
| 6/6 | ~/.agent-coordination/mail/.quarantine-test-leak-20260903/ |
12 leaked test messages moved here by exact id (0 remaining ×6) | they are the backup; nothing deleted |
| hub | ~/.agent-coordination/mail/.state/claude@rdmsm4x_55b61233.seen |
per-session seen-list, seeded from the address list (35,025 B) | additive |
| hub | ~/.agent-coordination/checkins/claude-rdmsm4x-feat-20260903-08.json |
check-in | additive |
| hub | ~/dev/fleet/SESSION-STATE.md |
§0 for this work, committed | git |
Repos
/ PRs (branches pushed to backup = private GitHub and
fleet = git.ecs0.net)
- richhdoty/rdmsm4x-dev-fleet-maintenance #3 —
scripts/agent_msg.zshv1.5.0,scripts/tests/test_agent_msg.zsh(47 checks),scripts/fleet_config_reconcile.zsh(distributes bus + hook hub→spokes, sha-verified, divergence-gated). Commitcc2f1f9preserves another session's uncommitted v1.4.0 verbatim;220e80eis v1.5.0. - richhdoty/rdmsm4x-dev-issues #1 —
bin/ticket: an address cannot hold a claim;handoff --to <address>records an OFFER; newaccept;claims.worktreecolumn (identity ↔︎ worktree ↔︎ claim);tests/test_offer_accept.py(44 checks).6365fe9. - richhdoty/rdmsm4x-dev-fleet #2 —
ops/git/hooks/prepare-commit-msgsession chain.6e3c196. - Isolated worktrees:
~/dev/_worktrees/fleet-maintenance-bus-v1.5.0,~/dev/_worktrees/issues-ticket-offers,~/dev/_worktrees/fleet-git-hook-session— keep until merged.
Tickets
FEAT-20260903-08 (this work, claimed) ·
FEAT-20260903-09 Tyrell chat/messaging spec, child of
FEAT-20260831-03, owner notified (20260903-153520-E22A54CB)
· ISSUE-20260903-29 spokes have no ticket
entrypoint (filed, not fixed).
Commands run (rerunnable)
zsh ~/dev/_worktrees/fleet-maintenance-bus-v1.5.0/scripts/tests/test_agent_msg.zsh # 47/47
BUS=~/.agent-coordination/agent_msg.zsh.bak-v1.4.0-<ts> zsh .../test_agent_msg.zsh # refused rc 2 (no NOSYNC gate)
cd ~/dev/_worktrees/issues-ticket-offers/tests && python3 test_offer_accept.py # 44/44
cd ~/dev/_worktrees/issues-ticket-offers/tests && python3 test_claim_leases_and_plans.py # 68/68
zsh ~/.agent-coordination/agent_msg.zsh whoami | roster --host rdmsm4x | inbox # live
# per spoke: scp + shasum compare (both files) + AGENT_SESSION=probe zsh agent_msg.zsh whoami
# on rdmbair15m5 (arm64) and rdmpw3265m (x86_64): the 47-check suite against the DEPLOYED file — 47/47 eachVerification evidence
- Live anycast round-trip on the real bus
20260903-153507-AD7922EA: two sessions seeNEW … [anycast];take→taken … by claude@rdmsm4x/55b61233; siblingtake→ rc 3 "already TAKEN … stand down"; sibling inboxtaken … [taken:55b61233], minemine; header carriesfrom_id:,session:,mode: anycast. Marked done. - Preflight (consumer of
inbox) still parses v1.5.0 output: rc 1 with the same 735 unacked HIGH the crowd had. - Roster on rdmsm4x: claude procs=6 registered=2 UNNAMED=4; codex procs=7 registered=0 UNNAMED=7 — the gap is printed, not hidden.
- Suite discriminates: against v1.4.0 it fails 20/47 (and is now refused outright).
- Seen-lists intact after the leak: 35,025 B on all six hosts, same mtime.
Incident (mine, contained)
Running the suite against the OLD v1.4.0 script spawned its un-gated
background sync, which rsynced 12 test messages from the
temp maildir to five spoke maildirs (two addressed to the real
claude@rdmsm4x). Quarantined on all six hosts by exact id;
suite now refuses any script lacking AGENT_MSG_NOSYNC.
Undo
Per host:
cp ~/.agent-coordination/agent_msg.zsh.bak-v1.4.0-<ts> ~/.agent-coordination/agent_msg.zsh
and the same for the hook .bak-<ts>;
mv ~/.agent-coordination/AGENT_COORDINATION.md.bak-pre-section9-<ts> ~/.agent-coordination/AGENT_COORDINATION.md
on the hub. Per-session seen files are additive and can be removed. PRs
can simply be closed.
Owner actions outstanding
- Merge PRs #3 (maintenance), #1 (issues), #2 (fleet). 2.
git -C ~/dev/issues pullon rdmsm4x — until then the ledger still accepts an address as holder. 3.git -C ~/dev/fleet/maintenance pullso bus_watch and the launchd reconciler run v1.5.0 + the distribution block. 4. Decide ISSUE-20260903-29 (spoketicketentrypoint).
codex_links ledger — PENDING (note not reachable from this Mac)
No Codex folder exists in any Notes account here and no
note named codex_links was found (21 codex-named notes
enumerated, none is the ledger). Filed as a fleet ISSUE with the five
verified rows; not re-created here to avoid a competing ledger. Rows:
PRs maintenance#3, issues#1, fleet#2 (private repos, OPEN via
gh); tickets FEAT-20260903-08 and -09 (behind Cloudflare
Access; origin 200).
Addendum 2026-09-03 20:35 EDT — Rich: "all approved" → merged and live
- Merged: rdmsm4x-dev-fleet-maintenance #3 →
d8aa69f(rebased onto the v1.4.0 author's own commite74d218, byte-identical to my preserved copy, so the duplicate dropped) and #4 →fea0f38(reconciler:typeseton an existing zsh parameter prints it — four straybf=lines per run; now plain assignments; proven un-quiet: 0 stray lines,agent_msg.zsh current+prepare-commit-msg currenton every reachable host). rdmsm4x-dev-issues #1 →212db7b, pulled on the authority host —ticket acceptlive,claims --jsoncarriesoffer. rdmsm4x-dev-fleet #2 →d911640; root main497d274, submodule pointerfea0f38, pushed to backup + fleet. - Broadcast of the approval and the new rules to
every host:
20260903-203237-34EA43CD. - Ticket FEAT-20260903-08 resolved, claim released, check-in withdrawn.
- jdmbair13m5 (corrected 20:40): bus + hook verified
deployed at 15:3x. Since ~20:3x the host refuses the hub's ed25519 key
on both
.localand.tspaths while up with a logged-in session and live agents (its 20:29 heartbeat lists claude/agy/codex processes) — so this is an auth-config change on that host, not the FileVault gate I first wrote here. FiledISSUE-20260903-36with thessh -vtrace and next steps for the agents on that host. The reconciler silently skips it until inbound ssh works; the deployed files were verified before the change. - Pre-existing defect noticed, not fixed:
~/dev/fleet/.gitmoduleshas no mapping for themaintenancegitlink (git submodule statuserrors); the pointer bumps still work via the index.
Addendum 2026-09-03 21:12 EDT — "fix all remaining issues", all approved
- ISSUE-20260903-29 resolved. Spoke
~/bin/ticketproxy v1.1 (a peer's hand-installed v1.0 + the policy §9 session chain; refuses to run on the authority), versioned asfleet/maintenance/scripts/ticket_proxy.zsh(PR #5 →5ed4666) and installed by a spokes-only reconciler block (absent / symlink / older proxy → install with.bak; foreign file → report). Deployed 4/4 (f35fd39a…), identity proven on each (agent@<spoke>/chain0k), negative claim refused with the holder named, write round-trip authoragent@rdmbair15m5/roundtrp2. Found and fixed on the way:ticket commentstamped the authority hostname as author → default is nowresolve_identity()(rdmsm4x-dev-issues PR #2 →88c84c5, pulled; suite 47/47). - ISSUE-20260903-32 resolved. No
codex_linksnote existed in any account on any reachable host (hub accounts enumerated; On My Mac accounts on both Intel hosts listed). Created in iCloud[email protected] › Codex › codex_linkswith the 5 verified rows; delivery verified on rdmbair15m5 (count 1); exact address + lookup snippet recorded in~/.claude/skills/project-link-ledger/SKILL.md. Owner action: pin it once (AppleScript cannot). - TASK-20260903-35 resolved.
~/Documents/Codexis a plain local folder mirrored hub→spokes; rdmsm4x holds the superset (29,598 files vs ≤29,011); spoke-only deltas are.DS_Storeonly; one same-size.git/indexspot-checked byte-identical. Nothing moved, nothing deleted. - Fleet root pointer
06c6a2c; a stale 0-byte.git/index.lock(4 min old, no git process) was removed first. - jdmbair13m5: unreachable (ISSUE-20260903-36). Two one-shot checks at 21:58 and 22:03 exist only in this session; the steps are in the ISSUE and SESSION-STATE for a hand run if the session is gone.
Addendum 2026-09-03 22:06 EDT — jdmbair13m5 recheck (Rich: "check again in an hour")
- Reachable again at 21:58:18
(
ssh -o BatchMode=yes … 'echo OK'→ OK, rc 0). Busagent_msg.zsh(5c2fdde422432e4c) and hookprepare-commit-msg(1ec70f1964bbc9d2) were still the deployed versions — nothing on the host had changed them. - Ticket proxy:
~/bin/ticketwas the peer's hand-installed v1.0 (217bf72e9bdc7cda) → replaced 21:59:50 with v1.1 (f35fd39abe5d045c, backup~/bin/ticket.bak-20260903-215950). Proof:AGENT_SESSION=probe ~/bin/ticket mine→Identity: agent@jdmbair13m5/probe, rc 0. All five spokes now run v1.1. - Reconciler pass 21:59:51
(
--host jdmbair13m5):agent_msg.zsh current,prepare-commit-msg current,ticket proxy current, settings current (defaultMode=bypassPermissionsuntouched), 8 skills synced (they had lagged while the host was unreachable — exactly the case the reconciler exists for).current=1 updated=0 diverged=0 offline=0 err=0. - ISSUE-20260903-36 resolved — cause not determinable from
retained logs, stated as such. Evidence:
authorized_keysunchanged (612 B, Aug 23, hub key present among 6),sshd_config.dunchanged (Feb),sshd -T= pubkey yes / strictmodes yes / LogLevel INFO; host never slept after 14:31; in the 20:31–20:36 refusal window 27 connections reached sshd and each ended insidesshd-authafter resolving user+groups, before the PAM/OpenDirectory account stage every accepted login runs — no StrictModes, sandbox, TCC or PAM denial logged. Failed publickey attempts are logged only atLogLevel VERBOSE, so the reason is not on disk. Recovered at 21:04:00 (Accepted publickey … uuN2C5…from the hub's LAN address), 53 min before Rich's 21:57 login, with no mtime change under~/.sshor/etc/ssh. Correction to my own filing: the reconciler does print<host>: offline — will retry next tickand countsoffline=N; "silently skips" was wrong. - Follow-up filed: fleet-wide sshd
LogLevel VERBOSEdrop-in so the next refusal names its reason (see ticket in SESSION-STATE §0). - Commented on ISSUE-20260903-29 that jdmbair13m5 is covered.