Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260829-2356-mcpreconcile-hardening-and-host-saturation

rdmsm4x changelog — 2026-08-29 23:49 → 23:57 EDT

Host: rdmsm4x · Session: claude@rdmsm4x · Trigger: cross-session relay from peer logtty-52 reporting a fleet resume and two fleet-sync breakages.


Peer claims — all three verified independently, none taken on trust

claim verdict
configsync was failing since 23:37; CLAUDE.md rewritten at 23:21 without the delegation heading TRUE. File is 17460 B, sha fb8fcc471e31c05f, mtime 23:21, and has exactly ONE ## heading.
codex/agy policy block frozen at 23:06 → two rulebooks TRUE, and see the open item below.
tyrell rebuilt arm64-only at 23:09, so mcpreconcile is correctly refusing TRUE — and it exposed a defect of mine.

Fixed — fleet_mcp_reconcile.zsh v1.0 → v1.2

Backup: fleet_mcp_reconcile.zsh.bak-20260829-2350.

v1.1 — the arch guard's blast radius was wrong, and that was my defect. v1.0 ran the universal-binary check as a global pre-flight that exit 1-ed on the first bad binary. So an arm64-only tyrell build stopped replicantdb reconciling on every host, for no reason. com.eastcoastscience.mcpreconcile was exit 1 from 23:33. Refusing to distribute to a mixed-arch fleet was right; halting unrelated servers was not. Guard is now per-server, and refusals surface in a new "refused" field in mcp-sync-state.json instead of looking like a dead job.

v1.2 — a concurrent rebuild makes healthy servers look dead. Fleet-wide hazard. At 23:50 the probe returned exit 137 (SIGKILL), no output for replicantdb on rdmsm4x. The server was fine: 8/8 clean probes minutes later, framing fix from 2026-08-26 still present in source. macOS SIGKILLs a running Mach-O whose file is overwritten — the binary being health-checked was the binary the peer's rebuild was writing. Because the reconciler's response to a failed handshake is to redeploy, a single-probe verdict would have pushed a mid-write binary across all six hosts. Now retries once after 3 s.

Generalise: anything that execs a binary and then repairs by copying it has this exposure while builds run — agentheal and fleetaudit are worth reviewing for single-probe verdicts.

Verified after: current=6 updated=0 offline=0 errors=0 refused=[], LaunchAgent exit 0, all six hosts answering a real pipelined initialize.

Confirmed working — the offline host self-healed

rdmpw3265m was unreachable for all of 2026-08-26 and received none of that session's repairs. The 23:03 run today shows it current, offline_pending: []. It converged with no hand-replay, which is what the reconciler was built to do.

Found while blocked — rdmsm4x is saturated

The ~/dev/todo intake scanner took over four minutes for a pass that normally takes about a second. The script is not at fault:

load averages: 75.26  100.55  89.77     (16 cores → 5-6x oversubscribed)

Largest single contributor: agy running rg nc -u -w 1 /Users/richh /etc — 13 minutes, 56% CPU, 1.8 GB RSS, searching the whole home directory and /etc for a literal string. Also 1417 processes (42 claude, 26 node), fileproviderd at 92% for 20 h, filecoordinationd at 49%. LoginItems.appex holds 2.5 GB RSS and is bursty, not pegged — one 100% sample and one 0.0% sample, so do not quote it as sustained.

Nothing was killed. The agy session holds live conversation state; automation heals infrastructure, never user work. Filed as [todo:20260829-rdmsm4x-saturated-load-100].

This is the most plausible explanation for the "performance has been sub-par" report of 2026-08-26 — a host-health problem, not model or config. Config was audited that day and clean.

Open for Rich — the CLAUDE.md rewrite dropped two standing rules

grep -c bypassPermissions ~/.claude/CLAUDE.md → 0. The response-timestamp rule is gone too. The setting is intact and ~/dev/fleet/CLAUDE.md rule 4 still forbids changing it, so the guardrail survives at domain level — but the global policy that made it fleet-wide does not. Combined with the frozen codex/agy block, Claude and codex/agy are on different rulebooks.

I did not edit CLAUDE.md and will not on a peer's say-so. Routed to Rich and to FLEET-STATUS-LEDGER.md.

Files touched

Both todo items are queued, not yet relayed — the scanner is crawling under the load above. The intake is idempotent and tag-keyed, so they land on the next completed pass without intervention.

Apple Notes entry pending — Notes is unresponsive fleet-wide per the standing ~/dev/ISSUES.md item; this file is the durable copy. Not silently skipped.


Addendum — 2026-08-30 00:00 → 00:09 EDT (peer exchange with logtty-52)

Four rounds with a peer session, each side finding defects in the other's area and in its own. Everything below was verified before being claimed.

Took the unowned probe-then-repair audit — both scripts clean

agent_heal.zsh and fleet_audit.zsh repair with launchctl kickstart -k only — no binary distribution — and both verify after and escalate if uncleared. The exposure needs exec-a-binary THEN copy-that-binary-on-failure; kickstart-only fails the second half. Worst case for a false positive is one unnecessary job restart. agent_heal reads launchd state rather than exec'ing anything, so the overwrite/SIGKILL trap cannot arise there at all.

fleet_mcp_reconcile.zsh v1.2 → v1.5

Also patched — two SHARED scripts, with the discipline that matters

fleet_config_reconcile.zsh (17 sites) and fleet_audit.zsh (3, not 2 — one line held two) got the same -k 5. Mechanical and additive only; no timeout value or logic changed. Confirmed no instance was running, published a check-in first (checkins/claude-rdmsm4x-20260830-000540-timeout-hardening.md), re-checked after. The peer and I edited fleet_config_reconcile.zsh a minute apart; my 17 sites survived intact and their complementary guard landed on top. Both halves were needed — hard timeouts made the failure prompt, their guard stopped it ending the run. Their real reconcile then completed in ~30 s, down from 8+ minutes, exit 0.

Corrections I made to my own earlier claims

Host health — unchanged and getting worse, not settling

Load re-spiked to 137 after falling to 51. fileproviderd ~100% for 20 hours is the durable tax; 8 concurrent swift/xcode processes are legitimate transient work. The agy rg nc -u -w 1 /Users/richh /etc was still running at 19 minutes, 2 GB RSS. Nothing killed.

Final state: mcpreconcile v1.5 live, LaunchAgent exit 0, current=6 refused=[] offline=[] errors=[]. Both queued todo items relayed into fleet/maintenance/ISSUES.md.


Addendum 2 — 2026-08-30 00:12 → 00:22 EDT

github_backup_dev.zsh v1.4 → v1.5 — four defects, one a safety-property bug

set -u, no -e, so it under-reports rather than aborting. GHUSER=$(gh api user) and TMPIDX=$(mktemp) were unguarded; push exit status was discarded into the log; the run always printed done and exited 0. The mktemp one is not a reporting bug — empty TMPIDX means GIT_INDEX_FILE="" and git falls back to the repo's real index, silently voiding the script's core promise of never touching worktree or index. v1.5 refuses on both, counts push outcomes, and exits non-zero on failure or on attempted > 0 && pushed == 0.

Framing correction, verified: the backup job is not broken and these fixes are preventive. backup-20260829-0330.log shows 89 repos processed, 40 snapshots pushed. Better, the log history closes the 2026-08-26 finding: scheduled -0330 logs exist for 08-27, 08-28, 08-29 and no date before. That absence proves it genuinely was not running; the three since prove the repair held. Sequence: broken 08-23→08-26, repaired 08-26, working nightly since.

fleet-config-sync.zsh v1.1.0 → v1.2.0 — a secret-exposure race, fixed on all six hosts

Fixed path ${TMPDIR:-/tmp}/fleetcfg-$H plus rm -rf "$STAGE" at start, no lock, hourly schedule. Two concurrent instances shared one directory and wiped each other's tree. $STAGE is also where credentials are scrubbed — an rm -f filename sweep and a grep -rlE secret-pattern sweep, both running after .claude/.codex/.gemini/memory/LaunchAgents are copied in — so a race in that window can publish an unscrubbed tree to a GitHub repo. The scrubbing was correct; its atomicity was not.

Fix: per-run mktemp -d, hard failure if empty, trap cleanup. Verified on the hub branch (514 files staged) and the spoke branch (493 files pushed), temp dirs cleaned both times. All six hosts were identical ab6a195ce1d606b2 before; all six verified 8aedb4e7eee6bf79 after. Check-in published first. A stale pre-fix staging directory holding 514 files was removed.

Left alone deliberately: line 19's [[ "$H" == "$HUB" ]] || true is a no-op, but the hub case is correctly handled further down — vestigial dead code, not a bug. Changing what reaches a public-facing repo is Rich's call.

Corrections I made to my own claims tonight

  1. fileproviderd is iCloud Drive, not Dropbox.
  2. The launchd configsync run had not aborted — my wait was on a stale PID. It completed, 00:10:57 … err=0. launchctl list's status is not evidence about the current run.
  3. My new BACKUP SUSPECT check fired on every dry run; excluded=0 came from a counter I never incremented.
  4. My first substitution grep missed unquoted probe=$(…).
  5. My check-in overstated the line-19 no-op.

Five self-corrections, all caught by testing rather than by review — which is the actual lesson of the night: a check that returns a confident answer without testing what it claims to test is the recurring defect, in shell scripts and in my own reasoning alike.