rdmsm4x — fleet sync restored; CLAUDE.md rewrite surfaced
Session: 2026-08-29 23:42 – 2026-08-30 00:05 EDT ·
rdmsm4x Projects:
~/dev/fleet/telemetry-logging,
~/dev/fleet/maintenance,
~/.agent-coordination
Continuation of the agent_session_logger.zsh work (see
rdmsm4x-changelog-20260829-2312-agent-session-logger-terminal-spam.md).
1. Terminal spam — completed, 5 of 5 hosts
rdmbair15m5 had been recorded as unreachable. It
was not down — its bare Tailscale name resolves to a
stale node (offline 8d); the live node is
rdmbair15m5-1 / 100.74.59.4. ssh
and ping by bare name both time out, which reads as a dead
host rather than as wrong-name resolution. Every fleet host has such a
twin (bare = stale, -1 = live); two of the live addresses
are the syslog collectors the logger already hardcodes.
v1.1.1 deployed to all 5 Macs, each identity-pinned by
scutil --get ComputerName as well as sha256
a5ae709b77c7222e, so a patch could not land on the wrong
machine. Real pty smoke test on rdmbair15m5: 0 job notices,
correct EXIT 0 / EXIT 1 / EXIT 0.
2. Fleet config sync was DOWN and distributing nothing
com.eastcoastscience.configsync had been failing every
30 min since 23:37. Cause:
~/.claude/CLAUDE.md was replaced at 23:21
without a ## Model, delegation & spend heading;
fleet_config_reconcile.zsh extracts that section to build
the Codex/agy policy block, and its sys.exit(...) under
set -eu aborted before the distribution
loop. A cosmetic sub-step took the whole fleet distribution
down.
Fixed: extraction is now non-fatal and loud — stderr
warning, POLICY-BLOCK-STALE in the log line,
policy_block_stale in the state file. It degrades to
last-known-good rather than silently freezing, which respects the
original author's explicit warning about frozen snapshots. Backup:
fleet_config_reconcile.zsh.bak-20260829-2345-prepolicyfallback.
Result: all 6 hosts on CLAUDE.md fb8fcc471e31c05f,
updated and verified,
defaultMode=bypassPermissions unchanged on every host. Job
log confirms current=6 err=0.
Still open: the Codex/agy policy block is frozen at 23:06. Claude harnesses and Codex/agy are on different rulebooks. Will not self-clear.
3. Fleet MCP sync was blocked
com.eastcoastscience.mcpreconcile was refusing to
distribute because a 23:09 build sweep left tyrell's
binaries arm64-only — a correct refusal for this
mixed-arch fleet. Rebuilt universal
(--arch arm64 --arch x86_64). That unblocked the probe
stage, which then exposed replicantdb-mcp answering
nothing. Handed to the two live replicantDB sessions rather than fixed
here; not reproducible by 23:56 (artifact replaced,
v1.10.1, all 6 hosts answer). Recorded as artifact changed,
not fixed — nobody can point to the
commit.
The same 23:09 sweep also left dist/replicantDB.app
unsigned with an empty Contents/MacOS/. That sweep
is the unexplained thing and is still unchased.
mcpreconcile now: exit 0, 6/6 current,
refused: [], errors: [] (peer
dev-ef made the arch guard per-server so one bad binary no
longer halts everything, and added a probe retry).
4. Headline for Rich — the 23:21 CLAUDE.md rewrite
198 lines → 99, entirely different content. 13 top-level
sections removed, including "The standing rule that
outranks the rest" — the rule that no agent may ever change
permissions.defaultMode. The live file now has
zero occurrences of defaultMode.
The setting itself is intact and correct
(bypassPermissions), verified. Only the written
guardrail is gone. No .bak was written at 23:21, unlike
every prior edit. A peer audit found no Rich fingerprint and a timing
match to an agy@rdmbair15m5 deploy.
Predecessor is safe in six places — preserved at
~/.agent-coordination/canonical/CLAUDE.md.PRE-REWRITE-20260829-2321.da3696457d339258.md,
and the same bytes survive as
CLAUDE.md.bak-reconcile-20260829-234728 on all five remote
hosts, written by the reconcile before it overwrote them.
Filed as ~/dev/DECISIONS-PENDING-RICH.md
§0 with three options and a one-line restore.
Not an agent decision — four sessions independently
declined to revert it on a peer's say-so.
5. Coordination
Fleet ledger at
~/.agent-coordination/FLEET-STATUS-LEDGER.md (append-only;
a peer session created its own version concurrently — I appended rather
than overwrote). All live peer sessions messaged with the resume
directive; broadcast sent to the cross-LLM bus for Codex/agy.
Corrections issued this session
- My
replicantdb-mcpdefect report → not reproducible, artifact changed under both of us. - A peer's "peer's 'repaired' claim false" on configsync →
corrected with the job's own log.
launchctl listis not a reliable health signal for a job you just kickstarted (-kSIGTERMs the instance, so the row shows-15); the job's log and state file are. This misled two sessions.
6. CORRECTION to §2 — the job was not yet reliable when I called it green
At 00:01 I reported config sync restored and green. Half
right. The 23:54:14 run did complete
(current=6 err=0) and the fleet was genuinely in sync — but
the next run aborted with exit 124 after 8+ minutes,
writing no log line and no state update. A run that aborts leaves no
trace, so nothing surfaced it until I checked the exit code
deliberately. Sync is in place and the job is reliable
are two claims and I ran them together.
Cause — the third instance tonight of a secondary
step killing the primary job. The codex/agy staging block was unguarded
under set -e:
STAGE="$(timeout 20 ssh ... 'mktemp -d -t cfgsync')" # command substitution, no || true
timeout 60 scp -q ... "$BLOCK" "$INJ" "$R:$STAGE/" # no || true
One slow host ended the whole reconcile, so every host after it was
silently never reconciled. X="$(timeout ...)" is the easy
one to miss — it reads as an assignment, not a command.
Fixed (wrapped in if/then/else, host skipped rather than
fatal) and proven: full real run at 00:06:37 →
exit 0, log line written, current=6 err=0,
POLICY-BLOCK-STALE still firing, runtime
~30s down from 8+ minutes. All 6 hosts re-verified on
fb8fcc471e31c05f by direct shasum, independent of the job's
own reporting — tonight exit codes and actual state disagreed in
both directions.
Probe traps worth keeping
- Overwriting a Mach-O SIGKILLs the running copy — a
health check during a build sees a healthy server as dead (exit 137).
mcpreconcileredeploys on failure, so a single-probe verdict during a build could have pushed a mid-write binary to all six hosts. - A missing
timeoutbinary on a remote host makes a probe exit 127 with empty stdout, indistinguishable from a server that answers nothing. A positive control must run on the same host and code path. - These two have one root: the apparatus that judges
a thing must be verified on the same host and code path as the thing it
judges. Concretely — every timeout in these scripts wraps ssh/scp
locally and must never move inside the ssh command, because
timeoutis Homebrew-provided here and absent from a non-login remote shell's PATH.
7. Swept the scheduled jobs for the same bug class — clean
Mapped all 15 com.eastcoastscience.* LaunchAgents plist
→ script (not all ~59 fleet scripts; the scheduled ones are where silent
failure costs something). The class splits on set -e:
- FATAL-ABORT (
set -e, aborts the run): onlyconfigsync(fixed, proven) andclaudestate.claudestate→claude_state_backup.zshchecked and clear — its two substitutions aredateandmktemp -d, local, no ssh/timeout. This mattered because a silent abort there stops the DR backup with no trace. - silent-empty
(
set -u/-uo): 13 jobs, 41+ substitution assignments.dev-effixed one real instance infleet_mcp_reconcile(a host reportingcurrentwhen its registration had silently failed) and showedfleet_auditis fail-loud by construction. I audited none of the rest — the count is a map, not a verdict.devbackup→github_backup_dev.zsh(8 subst,set -u) is the first I would look at: a backup that silently no-ops is the same shape.
Unowned / not claimed
— CLOSED byagenthealandfleetauditdev-ef: both clean. The exposure needs exec-a-binary then copy-it-on-failure; both repair withlaunchctl kickstart -konly and verify after. Their result, not re-verified by me.— RESOLVED.timeout Nnot a hard bounddev-efpatched all three scripts (29 call sites) totimeout -k 5 N, after proving it:timeout 2against a TERM-ignoring child ran the full 8s;timeout -k 1 2stopped at 3s.config reconcile aborts mid-run— RESOLVED and PROVEN. See §6 below.- The 23:09 build sweep itself.
- The silent-empty class across 13 scheduled jobs (above) — counted, not audited.
- A peer flagged that
ACTIVE-DIRECTIVES.mddropped three standing carve-outs — including delete-nothing on jdmbair13m5 — and Rich has not evidenced voiding them. Safety-relevant given the resume directive; I have not acted on it.