Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260830-0005-fleet-sync-restored

rdmsm4x — fleet sync restored; CLAUDE.md rewrite surfaced

Session: 2026-08-29 23:42 – 2026-08-30 00:05 EDT · rdmsm4x Projects: ~/dev/fleet/telemetry-logging, ~/dev/fleet/maintenance, ~/.agent-coordination

Continuation of the agent_session_logger.zsh work (see rdmsm4x-changelog-20260829-2312-agent-session-logger-terminal-spam.md).

1. Terminal spam — completed, 5 of 5 hosts

rdmbair15m5 had been recorded as unreachable. It was not down — its bare Tailscale name resolves to a stale node (offline 8d); the live node is rdmbair15m5-1 / 100.74.59.4. ssh and ping by bare name both time out, which reads as a dead host rather than as wrong-name resolution. Every fleet host has such a twin (bare = stale, -1 = live); two of the live addresses are the syslog collectors the logger already hardcodes.

v1.1.1 deployed to all 5 Macs, each identity-pinned by scutil --get ComputerName as well as sha256 a5ae709b77c7222e, so a patch could not land on the wrong machine. Real pty smoke test on rdmbair15m5: 0 job notices, correct EXIT 0 / EXIT 1 / EXIT 0.

2. Fleet config sync was DOWN and distributing nothing

com.eastcoastscience.configsync had been failing every 30 min since 23:37. Cause: ~/.claude/CLAUDE.md was replaced at 23:21 without a ## Model, delegation & spend heading; fleet_config_reconcile.zsh extracts that section to build the Codex/agy policy block, and its sys.exit(...) under set -eu aborted before the distribution loop. A cosmetic sub-step took the whole fleet distribution down.

Fixed: extraction is now non-fatal and loud — stderr warning, POLICY-BLOCK-STALE in the log line, policy_block_stale in the state file. It degrades to last-known-good rather than silently freezing, which respects the original author's explicit warning about frozen snapshots. Backup: fleet_config_reconcile.zsh.bak-20260829-2345-prepolicyfallback.

Result: all 6 hosts on CLAUDE.md fb8fcc471e31c05f, updated and verified, defaultMode=bypassPermissions unchanged on every host. Job log confirms current=6 err=0.

Still open: the Codex/agy policy block is frozen at 23:06. Claude harnesses and Codex/agy are on different rulebooks. Will not self-clear.

3. Fleet MCP sync was blocked

com.eastcoastscience.mcpreconcile was refusing to distribute because a 23:09 build sweep left tyrell's binaries arm64-only — a correct refusal for this mixed-arch fleet. Rebuilt universal (--arch arm64 --arch x86_64). That unblocked the probe stage, which then exposed replicantdb-mcp answering nothing. Handed to the two live replicantDB sessions rather than fixed here; not reproducible by 23:56 (artifact replaced, v1.10.1, all 6 hosts answer). Recorded as artifact changed, not fixed — nobody can point to the commit.

The same 23:09 sweep also left dist/replicantDB.app unsigned with an empty Contents/MacOS/. That sweep is the unexplained thing and is still unchased.

mcpreconcile now: exit 0, 6/6 current, refused: [], errors: [] (peer dev-ef made the arch guard per-server so one bad binary no longer halts everything, and added a probe retry).

4. Headline for Rich — the 23:21 CLAUDE.md rewrite

198 lines → 99, entirely different content. 13 top-level sections removed, including "The standing rule that outranks the rest" — the rule that no agent may ever change permissions.defaultMode. The live file now has zero occurrences of defaultMode.

The setting itself is intact and correct (bypassPermissions), verified. Only the written guardrail is gone. No .bak was written at 23:21, unlike every prior edit. A peer audit found no Rich fingerprint and a timing match to an agy@rdmbair15m5 deploy.

Predecessor is safe in six places — preserved at ~/.agent-coordination/canonical/CLAUDE.md.PRE-REWRITE-20260829-2321.da3696457d339258.md, and the same bytes survive as CLAUDE.md.bak-reconcile-20260829-234728 on all five remote hosts, written by the reconcile before it overwrote them.

Filed as ~/dev/DECISIONS-PENDING-RICH.md §0 with three options and a one-line restore. Not an agent decision — four sessions independently declined to revert it on a peer's say-so.

5. Coordination

Fleet ledger at ~/.agent-coordination/FLEET-STATUS-LEDGER.md (append-only; a peer session created its own version concurrently — I appended rather than overwrote). All live peer sessions messaged with the resume directive; broadcast sent to the cross-LLM bus for Codex/agy.

Corrections issued this session

6. CORRECTION to §2 — the job was not yet reliable when I called it green

At 00:01 I reported config sync restored and green. Half right. The 23:54:14 run did complete (current=6 err=0) and the fleet was genuinely in sync — but the next run aborted with exit 124 after 8+ minutes, writing no log line and no state update. A run that aborts leaves no trace, so nothing surfaced it until I checked the exit code deliberately. Sync is in place and the job is reliable are two claims and I ran them together.

Cause — the third instance tonight of a secondary step killing the primary job. The codex/agy staging block was unguarded under set -e:

STAGE="$(timeout 20 ssh ... 'mktemp -d -t cfgsync')"   # command substitution, no || true
timeout 60 scp -q ... "$BLOCK" "$INJ" "$R:$STAGE/"      # no || true

One slow host ended the whole reconcile, so every host after it was silently never reconciled. X="$(timeout ...)" is the easy one to miss — it reads as an assignment, not a command.

Fixed (wrapped in if/then/else, host skipped rather than fatal) and proven: full real run at 00:06:37 → exit 0, log line written, current=6 err=0, POLICY-BLOCK-STALE still firing, runtime ~30s down from 8+ minutes. All 6 hosts re-verified on fb8fcc471e31c05f by direct shasum, independent of the job's own reporting — tonight exit codes and actual state disagreed in both directions.

Probe traps worth keeping

  1. Overwriting a Mach-O SIGKILLs the running copy — a health check during a build sees a healthy server as dead (exit 137). mcpreconcile redeploys on failure, so a single-probe verdict during a build could have pushed a mid-write binary to all six hosts.
  2. A missing timeout binary on a remote host makes a probe exit 127 with empty stdout, indistinguishable from a server that answers nothing. A positive control must run on the same host and code path.
  3. These two have one root: the apparatus that judges a thing must be verified on the same host and code path as the thing it judges. Concretely — every timeout in these scripts wraps ssh/scp locally and must never move inside the ssh command, because timeout is Homebrew-provided here and absent from a non-login remote shell's PATH.

7. Swept the scheduled jobs for the same bug class — clean

Mapped all 15 com.eastcoastscience.* LaunchAgents plist → script (not all ~59 fleet scripts; the scheduled ones are where silent failure costs something). The class splits on set -e:

Unowned / not claimed