Fleet changelogs · dev.ecs0.net
rdmbair15m5-changelog-20260905-1151-herdr-full-mesh-and-session-resume

herdr full mesh + local session resume across the fleet

Session: claude@rdmbair15m5 (8e035a43) · Phase 2 2026-09-05 11:30 → 11:52 EDT Follows: rdmbair15m5-changelog-20260905-0510-herdr-fleet-rollout.md (install/service/integrations) Ask: "resume all local sessions, make sure all hosts attach to all other hosts via herdr"

Result — counts, not "green"

Measure Value
Mesh links (each host → 5 peers) 30 / 30
Local Claude sessions resumed 16 (4 × rdmsm4x, rdmbair15m5, rdmpw3265m, rdmpw3275m)
Agents natively detected, with state 16 — working / idle / blocked
Hosts with 0 resumed sessions 2 (rdmbair13m5, jdmbair13m5 — claude hangs, ISSUE-…-06)
Script sha parity across hosts 5/5 for fleet_herdr_resume.zsh, fleet_herdr_mesh.zsh

The design decision that made state work

Phase 1 established that herdr claims a pane for an agent by local process detection; over ssh the foreground process is ssh, so a remote agent reports no state. Rather than fight that, the mesh separates the two jobs:

Session discovery reads ~/.claude/projects/<encoded-cwd>/<uuid>.jsonl newest-first; the real working directory comes from the cwd field INSIDE the file, because the directory name is a lossy encoding (path separators and literal hyphens both become -) and must never be decoded back.

jdmbair13m5 could not ssh to rdmbair13m5/rdmbair15m5 at all. Tailscale carries duplicate nodes: the bare names are devices last seen 15 days ago, while the live Macs re-registered as rdmbair13m5-1 / rdmbair15m5-1. Local DNS separately returned stale LAN IPs on an unroutable subnet (192.168.1.48 vs the real 10.0.4.45). Added HostName …-1 aliases to that host's ~/.ssh/config — backed up, inserted BEFORE Host * because ssh takes the FIRST value per option, verified with ssh -G plus a real login. The stale tailnet nodes should be deleted at source.

Root cause narrowed: claude hangs on the two 13m5 Macs (HERDR-20260905-06)

sample of a hung process shows the whole call graph as _dyld_start + 0, footprint 112K — stuck at the dynamic linker's first instruction, before any Claude code runs. Not config, not auth: the 199 MB binary image never loads. Binary is byte-identical to working hosts; the difference is that working hosts carry com.apple.provenance on it and rdmbair13m5 does not. Correction: the earlier claim that killing 31 orphaned daemons "made it worse (112s → 146s)" is void — both figures were timeout values on a command that never completes. Two timeouts are not durations.

Three safety fixes made after they were caught

  1. fleet_herdr_resume.zsh --close matched every claude@* label, which included the workspace running the script. Now protected via HERDR_WORKSPACE_ID; verified both ways — it closed 4 workspaces and spared w2.
  2. Sessions with no recoverable cwd used to fall back to $HOME, which makes Claude raise a trust prompt for the entire home directory. They are now skipped — that is a security decision for a human, not a default.
  3. Labels are made unique by probing the live workspace list rather than assuming an index (a duplicate claude@rdmbair15m5 had already appeared).

Also recorded: killing processes matched by ps | awk /claude --version/ killed the ssh session running the command, because that session's own argv contains the string. The [c] bracket trick guards the pattern, not the caller.

Left for Rich (deliberately not auto-answered)

Artifacts: rdmsm4x:~/dev/fleet/herdr/ — fleet_herdr_mesh.zsh, fleet_herdr_resume.zsh, fleet_herdr_attach.zsh, fleet_herdr_setup.zsh, herdr_tcc_grant.zsh, config.toml, ISSUES.md (5 OPEN, 3 RESOLVED), SESSION-STATE.md. No secrets written to any file, note or log.