rdmsm4x changelog — 2026-08-29 23:49 → 23:57 EDT
Host: rdmsm4x ·
Session: claude@rdmsm4x · Trigger:
cross-session relay from peer logtty-52 reporting a fleet
resume and two fleet-sync breakages.
Peer claims — all three verified independently, none taken on trust
| claim | verdict |
|---|---|
| configsync was failing since 23:37; CLAUDE.md rewritten at 23:21 without the delegation heading | TRUE. File is 17460 B, sha
fb8fcc471e31c05f, mtime 23:21, and has exactly ONE
## heading. |
| codex/agy policy block frozen at 23:06 → two rulebooks | TRUE, and see the open item below. |
| tyrell rebuilt arm64-only at 23:09, so mcpreconcile is correctly refusing | TRUE — and it exposed a defect of mine. |
Fixed —
fleet_mcp_reconcile.zsh v1.0 → v1.2
Backup: fleet_mcp_reconcile.zsh.bak-20260829-2350.
v1.1 — the arch guard's blast radius was wrong, and that was
my defect. v1.0 ran the universal-binary check as a
global pre-flight that exit 1-ed on the
first bad binary. So an arm64-only tyrell build stopped
replicantdb reconciling on every host, for no reason.
com.eastcoastscience.mcpreconcile was exit 1 from 23:33.
Refusing to distribute to a mixed-arch fleet was right; halting
unrelated servers was not. Guard is now per-server, and refusals surface
in a new "refused" field in
mcp-sync-state.json instead of looking like a dead job.
v1.2 — a concurrent rebuild makes healthy servers look dead. Fleet-wide hazard. At 23:50 the probe returned exit 137 (SIGKILL), no output for replicantdb on rdmsm4x. The server was fine: 8/8 clean probes minutes later, framing fix from 2026-08-26 still present in source. macOS SIGKILLs a running Mach-O whose file is overwritten — the binary being health-checked was the binary the peer's rebuild was writing. Because the reconciler's response to a failed handshake is to redeploy, a single-probe verdict would have pushed a mid-write binary across all six hosts. Now retries once after 3 s.
Generalise: anything that execs a binary and then repairs by
copying it has this exposure while builds run — agentheal
and fleetaudit are worth reviewing for single-probe
verdicts.
Verified after:
current=6 updated=0 offline=0 errors=0 refused=[],
LaunchAgent exit 0, all six hosts answering a real pipelined
initialize.
Confirmed working — the offline host self-healed
rdmpw3265m was unreachable for all of 2026-08-26 and
received none of that session's repairs. The 23:03 run today shows it
current, offline_pending: []. It converged
with no hand-replay, which is what the reconciler was
built to do.
Found while blocked — rdmsm4x is saturated
The ~/dev/todo intake scanner took over four
minutes for a pass that normally takes about a second. The
script is not at fault:
load averages: 75.26 100.55 89.77 (16 cores → 5-6x oversubscribed)
Largest single contributor: agy running
rg nc -u -w 1 /Users/richh /etc — 13 minutes, 56%
CPU, 1.8 GB RSS, searching the whole home directory and
/etc for a literal string. Also 1417 processes (42
claude, 26 node), fileproviderd
at 92% for 20 h, filecoordinationd at 49%.
LoginItems.appex holds 2.5 GB RSS and is bursty, not pegged
— one 100% sample and one 0.0% sample, so do not quote it as
sustained.
Nothing was killed. The agy session holds live
conversation state; automation heals infrastructure, never user work.
Filed as [todo:20260829-rdmsm4x-saturated-load-100].
This is the most plausible explanation for the "performance has been sub-par" report of 2026-08-26 — a host-health problem, not model or config. Config was audited that day and clean.
Open for Rich — the CLAUDE.md rewrite dropped two standing rules
grep -c bypassPermissions ~/.claude/CLAUDE.md →
0. The response-timestamp rule is gone too. The
setting is intact and ~/dev/fleet/CLAUDE.md rule 4
still forbids changing it, so the guardrail survives at domain level —
but the global policy that made it fleet-wide does not. Combined with
the frozen codex/agy block, Claude and codex/agy are on different
rulebooks.
I did not edit CLAUDE.md and will not on a peer's
say-so. Routed to Rich and to
FLEET-STATUS-LEDGER.md.
Files touched
~/dev/fleet/maintenance/scripts/fleet_mcp_reconcile.zsh(v1.0 → v1.2) +.bak-20260829-2350~/.agent-coordination/FLEET-STATUS-LEDGER.md(appended)~/dev/todo/items/20260829-mcpreconcile-blast-radius-and-rebuild-hazard.md(new)~/dev/todo/items/20260829-rdmsm4x-saturated-load-100.md(new)
Both todo items are queued, not yet relayed — the scanner is crawling under the load above. The intake is idempotent and tag-keyed, so they land on the next completed pass without intervention.
Apple Notes entry pending — Notes is unresponsive
fleet-wide per the standing ~/dev/ISSUES.md item; this file
is the durable copy. Not silently skipped.
Addendum
— 2026-08-30 00:00 → 00:09 EDT (peer exchange with
logtty-52)
Four rounds with a peer session, each side finding defects in the other's area and in its own. Everything below was verified before being claimed.
Took the unowned probe-then-repair audit — both scripts clean
agent_heal.zsh and fleet_audit.zsh repair
with launchctl kickstart -k only — no
binary distribution — and both verify after and escalate if uncleared.
The exposure needs exec-a-binary THEN
copy-that-binary-on-failure; kickstart-only fails the second half.
Worst case for a false positive is one unnecessary job restart.
agent_heal reads launchd state rather than exec'ing
anything, so the overwrite/SIGKILL trap cannot arise there at all.
fleet_mcp_reconcile.zsh
v1.2 → v1.5
- v1.3 — positive control. Empty stdout has many
causes that are not a failing server (SIGKILL mid-overwrite, a helper
missing from the remote PATH, ssh hiccup), and this script
redeploys on that verdict. It now re-runs the same delivery
path on the same host with the server swapped for
cat; no control, no verdict. Tested three ways including a regression check that a genuinely dead server still repairs. - v1.4 —
timeout -k 5at all 9 sites. Verified the premise first rather than inferring it:timeout 2against a TERM-ignoring child ran the full 8.023 s;timeout -k 1 2died at 3.024 s. Baretimeoutis not a bound. - v1.5 — positive confirmation from the registration
step. Line 247 tested
[[ "$OUT" == *updated* ]], and both branches end in|| true, so a step that died outright leftOUTempty — not*updated*, so it fell through and the host was recordedcurrent. A host whose registration silently failed reported healthy. Now requiresclaude:/codex:in the output.
Also patched — two SHARED scripts, with the discipline that matters
fleet_config_reconcile.zsh (17 sites) and
fleet_audit.zsh (3, not 2 — one line held two) got the same
-k 5. Mechanical and additive only; no timeout value or
logic changed. Confirmed no instance was running, published a
check-in first
(checkins/claude-rdmsm4x-20260830-000540-timeout-hardening.md),
re-checked after. The peer and I edited
fleet_config_reconcile.zsh a minute apart; my 17 sites
survived intact and their complementary guard landed on top. Both halves
were needed — hard timeouts made the failure prompt, their guard stopped
it ending the run. Their real reconcile then completed in ~30
s, down from 8+ minutes, exit 0.
Corrections I made to my own earlier claims
- The
fileproviderdstorm is iCloud Drive (com.apple.CloudDocs.iCloudDriveFileProviderManaged), not Dropbox as I first guessed. - A
2>&1 >/dev/nullredirection briefly made me think responses were on stderr. They are not. rdmbair13m5.local→192.168.0.131looked like a stale-address hazard against Tailscale's192.168.1.177. It is the reverse:.localanswers, the Tailscale endpoint is stale. No finding.- My first substitution grep (
="\$() missedfleet_audit's unquotedprobe=$(…)and would have let me call it clean for the wrong reason. Correct pattern:=\$(\|="\$(.
Host health — unchanged and getting worse, not settling
Load re-spiked to 137 after falling to 51.
fileproviderd ~100% for 20 hours is the
durable tax; 8 concurrent swift/xcode processes are legitimate transient
work. The agy rg nc -u -w 1 /Users/richh /etc
was still running at 19 minutes, 2 GB RSS. Nothing
killed.
Final state: mcpreconcile v1.5 live, LaunchAgent
exit 0, current=6 refused=[] offline=[] errors=[]. Both
queued todo items relayed into
fleet/maintenance/ISSUES.md.
Addendum 2 — 2026-08-30 00:12 → 00:22 EDT
github_backup_dev.zsh
v1.4 → v1.5 — four defects, one a safety-property bug
set -u, no -e, so it under-reports rather
than aborting. GHUSER=$(gh api user) and
TMPIDX=$(mktemp) were unguarded; push exit status was
discarded into the log; the run always printed done and
exited 0. The mktemp one is not a reporting bug — empty
TMPIDX means GIT_INDEX_FILE="" and git falls
back to the repo's real index, silently voiding the
script's core promise of never touching worktree or index. v1.5 refuses
on both, counts push outcomes, and exits non-zero on failure or on
attempted > 0 && pushed == 0.
Framing correction, verified: the backup job is
not broken and these fixes are preventive.
backup-20260829-0330.log shows 89 repos processed,
40 snapshots pushed. Better, the log history closes the
2026-08-26 finding: scheduled -0330 logs exist for
08-27, 08-28, 08-29 and no date before. That absence
proves it genuinely was not running; the three since prove the repair
held. Sequence: broken 08-23→08-26, repaired 08-26, working nightly
since.
fleet-config-sync.zsh
v1.1.0 → v1.2.0 — a secret-exposure race, fixed on all six hosts
Fixed path ${TMPDIR:-/tmp}/fleetcfg-$H plus
rm -rf "$STAGE" at start, no lock, hourly schedule. Two
concurrent instances shared one directory and wiped each other's tree.
$STAGE is also where credentials are
scrubbed — an rm -f filename sweep and a
grep -rlE secret-pattern sweep, both running
after
.claude/.codex/.gemini/memory/LaunchAgents
are copied in — so a race in that window can publish an
unscrubbed tree to a GitHub repo. The scrubbing was
correct; its atomicity was not.
Fix: per-run mktemp -d, hard failure if empty,
trap cleanup. Verified on the hub branch (514 files staged)
and the spoke branch (493 files pushed), temp dirs cleaned both
times. All six hosts were identical ab6a195ce1d606b2
before; all six verified 8aedb4e7eee6bf79 after. Check-in
published first. A stale pre-fix staging directory holding 514 files was
removed.
Left alone deliberately: line 19's
[[ "$H" == "$HUB" ]] || true is a no-op, but the hub case
is correctly handled further down — vestigial dead code, not a bug.
Changing what reaches a public-facing repo is Rich's call.
Corrections I made to my own claims tonight
fileproviderdis iCloud Drive, not Dropbox.- The launchd
configsyncrun had not aborted — my wait was on a stale PID. It completed,00:10:57 … err=0.launchctl list's status is not evidence about the current run. - My new
BACKUP SUSPECTcheck fired on every dry run;excluded=0came from a counter I never incremented. - My first substitution grep missed unquoted
probe=$(…). - My check-in overstated the line-19 no-op.
Five self-corrections, all caught by testing rather than by review — which is the actual lesson of the night: a check that returns a confident answer without testing what it claims to test is the recurring defect, in shell scripts and in my own reasoning alike.