Fleet changelogs · dev.ecs0.net
rdmbair15m5-changelog-20260823-0218-two-trap-repair-devauth-guarded-updater

Repairing the two traps, standing dev authorization, and a guarded Claude Code updater

Host: rdmbair15m5 · Session: Claude Code richh-69 · Window: 2026-08-23 02:12 → 02:18 EDT

Trap 1 — scp + mode 555 + && silently skipped three hosts. FIXED.

Two compounding faults: the canonical Claude statusline is installed mode 555, so scp fails with dest open … Permission denied; and my loop chained scp … && ssh install, so the short-circuit skipped the install entirely on three hosts while the loop reported nothing.

Repaired by building propagation into the tool as statusline_sync.zsh --push, which:

Verified: ALL HOSTS CONSISTENT, exit 0, 8/8 files on all five hosts, all three CLIs.

Trap 2 — the bare hostname resolves to a DEAD node. Worked around; root fix needs Rich.

This is not specific to rdmbair13m5. Every Mac in the fleet has two tailnet nodes:

Bare name (MagicDNS answers here) Live node
rdmsm4x 100.97.225.81 — offline since 2026-08-21 08:44 rdmsm4x-1 100.87.66.12
rdmbair13m5 100.106.221.104 — offline since 2026-08-21 08:40 rdmbair13m5-1 100.75.186.59
rdmbair15m5 100.75.253.27 — offline since 2026-08-21 08:40 rdmbair15m5-1 100.74.59.4
rdmpw3265m 100.102.96.86 — offline since 2026-08-21 08:40 rdmpw3265m-1 100.113.170.117
rdmpw3275m 100.72.32.84 — offline since 2026-08-21 08:38 rdmpw3275m-1 100.95.173.70

dscacheutil -q host -a name rdmbair13m5 returns 100.106.221.104 — the dead twin. So the bare name has been silently wrong fleet-wide for two days, and .local only works while the Mac is awake.

Workaround, deployed to all five hosts: ~/.agent-coordination/fleet_host.zsh

fleet_host.zsh resolve <host>   # tries <host>.local, then <host>-1, then bare; caches the winner
fleet_host.zsh ssh <host> 'cmd'
fleet_host.zsh check            # reachability table

Order matters: the bare name is tried last because it is the dead one. rdmbair13m5 now resolves to rdmbair13m5-1 automatically, and every propagation in this session used it.

Root fix — needs Rich, I have no Tailscale admin credential (nothing in ~/.secrets/global.env or Keychain, and tailscale set --hostname only renames the local node):

  1. https://login.tailscale.com/admin/machines
  2. Delete the five nodes listed in the left column above — all offline since 2026-08-21, all duplicates.
  3. Rename each -1 node to the bare name.

Reversible alternative if deleting feels heavy: rename the stale ones to <name>-old first, then rename the live ones — no node is destroyed. If Rich would rather I did it, put a Tailscale API key in ~/.secrets/global.env as TAILSCALE_API_KEY and I can do it through the API.

Standing development authorization — applied fleet-wide

Rich: "pre-authorize and/or fully authorize development work to skip prompts where possible, all development is authorized going forward."

Appended a FLEET-DEV-AUTHORIZATION v1.0 block to ~/.claude/CLAUDE.md, ~/.gemini/GEMINI.md and ~/.codex/AGENTS.md — 15 files across 5 hosts, each backed up first. Pre-authorized: build, test, lint, format, benchmark, profile; file edits inside the project; local git; Homebrew; local dev servers and launchd agents you own; diagnostic capture; reaping your own and hung test processes.

Explicitly not authorized: destroying data that is not yours, --force push or shared main, killing an agent session holding unsaved context, outward-facing actions, spending, weakening security, and Tailscale/identity changes. The test: reversible and confined to the project → do it; destroys something or leaves the machine → ask.

Harness state measured rather than assumed:

rdmbair15m5  claude=bypassPermissions  agy=always-proceed  codex=never/danger-full-access
rdmbair13m5  claude=bypassPermissions  agy=always-proceed  codex=never/danger-full-access
rdmpw3265m   claude=bypassPermissions  agy=always-proceed  codex=never/danger-full-access
rdmpw3275m   claude=bypassPermissions  agy=always-proceed  codex=never/danger-full-access
rdmsm4x      claude=auto  <-- only outlier

I deliberately did not script rdmsm4x. permissions.defaultMode is owned by the live session, which rewrites the file within minutes — documented in CLAUDE.md and observed fleet-wide on 2026-08-20. Asked the rdmsm4x session to set it in-session, or to confirm auto is intentional there (it is the only host that can sign releases, so a conservative default is defensible).

Claude Code update — already current, procedure built for next time

All five hosts: installed 2.1.241, latest published 2.1.241. Nothing to update.

Built ~/.claude/canonical/claude_update_safe.zsh (deployed to all five) which refuses to update unless it is safe:

  1. Compares installed vs the cask API; exits 0 if current; aborts rather than guessing if the API is unreachable.
  2. Inspects every agy tab via Terminal — busy, or any subagents running, blocks with exit 10.
  3. Checkpoints before anything restarts: prompts each agy session to write full state to SESSION-STATE.md + ISSUES.md and sync the bus, then waits 90s.
  4. Only then brew upgrade --cask claude-code@latest, and verifies the resulting version.

--check reports without acting; --force skips the idle gate but still checkpoints. Running sessions keep the old binary until they restart, so an update is never disruptive mid-task.

Files added this session

~/.agent-coordination/fleet_host.zsh              (all 5 hosts)
~/.claude/canonical/claude_update_safe.zsh        (all 5 hosts)
~/.claude/canonical/statusline_sync.zsh --push    (all 5 hosts)