Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260825-1615-fleet-config-reconciler-offline-catchup-and-devsite-documentation

Fleet config reconciler — offline hosts self-heal; thread documented to dev.dataroo.net

Host: rdmsm4x (lead) · Session: dev-db [982a7e] · 2026-08-25, this segment 16:05 → 16:16 EDT America/New_York Thread total: 13:32 → 16:16 EDT (CLAUDE.md v2.8 → v3.1, plus the delivery mechanism below)

Closes the gap the policy work exposed: a Mac that is asleep when config changes are pushed had no way to receive them. It now converges automatically. Also documents the whole thread to the dev wiki.

The problem

jdmbair13m5 was unreachable for this entire session, so it missed both v3.0 and v3.1. The existing tooling was one-shot: deploy_model_policy_v3.zsh pushed once, reported the host unreachable, and that was the end of it. Catching up required a human to remember and re-run it. Nothing on any host watched for a returning peer.

What was built

fleet_config_reconcile.zsh (v1.0) — convergence, not a delta queue

~/dev/fleet/maintenance/scripts/fleet_config_reconcile.zsh. Version-neutral, idempotent, re-runnable. Supersedes the one-shot deploy_model_policy_v3.zsh (kept for changelog references).

Deliberately a reconciler rather than a queue: the desired state is "every host matches canonical", so each tick simply compares and converges. There is no queue to drain, nothing to replay by hand, and no way for a queued delta to be lost or applied twice. A host that is current costs one SSH and prints current.

Reconciles ~/.claude/CLAUDE.md (whole file, gated), ~/.codex/AGENTS.md and ~/.gemini/GEMINI.md (delimited block only, so peer-written content outside the markers survives byte for byte).

com.eastcoastscience.configsync — launchd, every 30 min + RunAtLoad

Installed and bootstrapped on rdmsm4x. First run verified: exit code 0, empty stderr, log line written. State at ~/.agent-coordination/config-sync-state.json, rolling log at ~/.agent-coordination/config-sync.log (bounded to 500 lines).

{"last_run":"2026-08-25 16:15:17 EDT","canonical_sha":"1306d4dd6406e707",
 "current":["rdmsm4x","rdmbair13m5","rdmbair15m5","rdmpw3265m","rdmpw3275m"],
 "updated":[],"diverged":[],"offline_pending":["jdmbair13m5"],"errors":[]}

offline_pending is the answer to "queued changes": jdmbair13m5 is tracked, and the first tick after it wakes delivers CLAUDE.md and both harness blocks with no human involved.

Scope limit, stated rather than implied: the reconciler runs only on rdmsm4x. If rdmsm4x itself is offline, peers do not converge — acceptable, because rdmsm4x is the canonical source and changes originate there, so there is nothing to converge toward while it is down.

A defect the test caught — the mechanism would have reported clean and healed nothing

Rather than assume the catch-up worked, it was tested by deliberately reverting a live host (rdmpw3275m) to the old v2.8 policy and running the reconciler. It refused to heal it, reporting 113 diverged lines.

Root cause: the divergence gate compared a peer against only the current canonical. That cannot distinguish stale from locally edited. A host offline across a consolidation legitimately holds every line of the OLD canonical, so it scores as heavily diverged — meaning every returning offline host, the exact population the mechanism exists to serve, would have been blocked from healing permanently while the summary line still read clean.

Fixed with _divergence_check.py: a peer line counts as local only if it appears in no canonical version ever published (current + .last-deployed + every CLAUDE.md.bak*). Everything else is history, and history is safe to replace.

Both halves then proven on a live host, not asserted

Scenario Expected Observed
Host reverted to v2.8 (stale) heal to v3.1 CLAUDE.md updated + verified, header v3.1, sha 1306d4dd6406e707, blast-radius table present
Host carrying a novel local line refuse, preserve DIVERGED (1 line(s) in no known canonical) — not overwriting; the local line survived intact
After cleanup back to current current=1 updated=0 diverged=0

Second drift risk closed

The reconciler initially pushed a stored snapshot of the policy block to codex/agy. That block would have silently frozen at the day it was copied while CLAUDE.md moved on. It now regenerates the block from the canonical CLAUDE.md on every run by extracting the policy section directly. Verified by deleting the block file and confirming the next run rebuilt it with current v3.1 content. The orientation header was also made version-neutral so it cannot assert a stale version number.

Documented to dev.dataroo.net

Wrote ~/dev/data/status-updates/fleet.json — the documented intake path (merge_status.py folds agent reports into data/products.json; separate files per agent avoid a lost-update race). Merged and regenerated: 17 product pages, ~/dataroo.net/wiki/p/fleet.html rebuilt at 16:14.

Content verified present in the generated page: v3.1, blast radius, 1306d4dd6406e707, offline_pending, fleet_config_reconcile, 98.6.

Carried forward rather than wiped: the still-open Tailscale duplicate-node blocker from 2026-08-22 (merge_status.py replaces a product's lists wholesale, so an unrelated open item would otherwise have been destroyed by this report).

Honest limit — the public page was NOT verified. https://dev.dataroo.net/ returns HTTP 401 with www-authenticate: Basic realm="dataroo development wiki" and a purpose-built "Sign-in required" page. That is intentional access control, not a fault. The credential is not in ~/.secrets/global.env, so verification stopped at the generated file. Logged OPEN.

Reproduce / undo

zsh ~/dev/fleet/maintenance/scripts/fleet_config_reconcile.zsh --dry-run
cat ~/.agent-coordination/config-sync-state.json
tail -5 ~/.agent-coordination/config-sync.log
launchctl print gui/$(id -u)/com.eastcoastscience.configsync | grep -E 'runs =|last exit'
# disable the automation:
launchctl bootout gui/$(id -u)/com.eastcoastscience.configsync
# per-host undo (written before every write):
#   ~/.claude/CLAUDE.md.bak-reconcile-<ts>

Outstanding

  1. jdmbair13m5 — queued, not yet delivered. Self-heals on its first tick after waking; the automatic path is proven only by proxy (rdmpw3275m), never against this host.
  2. dev wiki basic-auth credential missing from ~/.secrets/global.env — blocks live page verification by any agent.
  3. Provisional usage ceilings; Fable re-measure ~2026-09-01 (carried from the v3.1 record).

Budget note: usage_status read advice: hold at 15:21 EDT (session_5h 98.6% of a provisional ceiling). Honored — this entire segment ran inline: zero subagents, zero codex, zero agy.

No secrets were written to any file, message, or note.