Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260831-1245-reconciler-skill-derivation-and-dryrun-visibility

rdmsm4x-changelog-20260831-1245-reconciler-skill-derivation-and-dryrun-visibility

Host: rdmsm4x · Session: dev-73 (claude@rdmsm4x) Window: 2026-08-31 12:42:48 → 12:45 EDT Scope: ~/dev/fleet/maintenance/ (reconciler + new git repo), verification across all 6 hosts.

Closed the ISSUE-20260830-42 residual jointly with dev-8a, verified their fix independently, and closed both follow-ups they raised. The residual I had filed as latent turned out to be already live — a second skill was missing fleet-wide.


1. Verified dev-8a's fix rather than relaying it

dev-8a replaced the hand-maintained skill list with derivation by inclusion: a local skill dir ships only if its name appears in ~/.claude/CLAUDE.md or ~/dev/CLAUDE.md (fleet_config_reconcile.zsh:53-60).

Checked myself:

Claim Result
Resolves to exactly the referenced skills 7 — claude-state-backup, fleet-notes-publish, fleet-operations, fleet-ticketing, macos-log-triage, project-link-ledger, terminal-shell-env
No overshoot 31 local dirs, 24 correctly excluded
terminal-shell-env now on all 5 peers sha fc5be8d5… identical on canonical + all 5

The prediction was right and the instance already existed. terminal-shell-env was absent on all five peers while CLAUDE.md's load-on-demand table pointed straight at it. A dead pointer fails silently — the skill simply is not there when an agent reaches for it.

A sha discrepancy I explained wrongly — corrected

dev-8a reported sha 30986f5f7fe60102; I measured fc5be8d58541a0f2. Both numbers were correct and both were internally consistent across all six hosts. I attributed the gap to "the file changed between their check and mine". That event never happened.

dev-8a pushed back and was right. Verified myself:

This is the session's own defect class arriving by another route: a real measurement acquiring a conclusion it does not support — the same shape as the "268 pages live on the site" alarm. Leaving it in the record would have sent someone hunting a writer that does not exist, which is precisely the fifteen hours two other sessions just lost hunting an owner for work that had already been reported.

Rule this reinforces: before narrating a difference between two measurements as an event, confirm the two measurements are the same function — and check mtimes first, since they falsify a phantom write immediately and cost nothing.

2. Closed follow-up #1 — --dry-run could not see the gap it existed to find

The entire skill-sync block sat inside if (( ! DRY )). So --dry-run printed current for every host while a referenced skill was missing on all of them. The one command an operator would run to check the fleet was structurally incapable of seeing this class of gap — which is a large part of why it survived a week.

Dry run now probes each peer per skill:

rdmbair13m5: would sync skill terminal-shell-env (present)
rdmbair13m5: would sync skill <name> — MISSING ON PEER

Verified both directions, per the standard added yesterday:

3. Closed follow-up #2 — the reconciler had no version history

fleet_config_reconcile.zsh writes CLAUDE.md, AGENTS.md, GEMINI.md, settings.json keys and 7 skills to all six hosts on a 30-minute timer, and had only sibling .bak-* files standing in for version control. Same gap class as the ticket store on 2026-08-30, higher blast radius.

~/dev/fleet/maintenance/ is now a git repo — bc38816, 41 files tracked, 376K of history.

Scoped there deliberately. ~/dev/fleet is 10G across 22,718 files, with 9.1G in ops/ and four nested git repos (macos, portfolio-consolidation, rdm-fleet, coordination-agy). A repo at the domain root would have been wrong. Ignores staging/ (2.7M transient), backups/, and the .bak-* siblings.

I amended my own baseline commit message. It said "captured as found" while actually containing the dry-run edit made moments earlier. Left uncorrected, that change would read as a pre-existing condition in the first diff anyone ran.

4. Verification

derivation           7 skills / 31 local dirs, 24 excluded
dry-run coverage     42 lines = 7 skills x 6 hosts
full reconcile       current=6 updated=0 diverged=0 offline=0 err=0
settings             defaultMode=bypassPermissions on all 6 — verified, NOT modified
terminal-shell-env   fc5be8d5 on canonical + all 5 peers

All peer checks used richh@<host>.ts.dataroo.net — the same address form the reconciler itself uses, per the rule added yesterday after a bare hostname nearly produced a false failure report.

5. A detail worth Rich's eye

dev-8a's own verification command failed with a bad pattern error on *(N/), because Rich's interactive zsh disables bare glob qualifiers — a trap documented in terminal-shell-env, the exact skill that was missing from every peer. The missing documentation caused the error that the documentation would have prevented.

The same trap hit this session last night: a ~/Desktop/*.html(N) count silently returned 0, and it was caught only by recounting with /bin/ls. Scripts are unaffected — they do not source the interactive rc — so this bites agents running ad-hoc commands, which is most of what we do.

Outstanding

  1. ~/dev/fleet/maintenance has no remote — history, not off-host backup. Same shape as the issue store, which dev-ec is closing right now via a private richhdoty/rdmsm4x-dev-issues.
  2. ~/dev/fleet at large is still unversioned, including ops/ at 9.1G. Not obviously wrong — much of it is output, not source — but nobody has decided which parts deserve history.
  3. The standards added yesterday are enforced by documentation, not tooling. A pre-commit hook rejecting git add -A in shared trees remains the highest-value mechanical enforcement.

Not touched

~/dev/issues beyond resolving ISSUE-20260830-42 — dev-ec owns the remote work there. No secrets read, written, or referenced. Nothing pushed to any remote. permissions.defaultMode verified bypassPermissions on all six and not modified.