herdr — full mesh, fleet-wide session resume, and two host repairs
Session: claude@rdmbair15m5 (8e035a43) ·
Phase 3 2026-09-05 11:53 → 12:30 EDT
Follows: …-0510-herdr-fleet-rollout.md,
…-1151-herdr-full-mesh-and-session-resume.md
Ask: "make all best judgements for me" / "all should
resume accordingly"
Result
| Measure | Value |
|---|---|
| Mesh links | 25/30 verified (rdmsm4x's own 5 verified 12:24, unqueryable since) |
| Agents fleet-wide | 33, 0 blocked |
| Hosts running resumed sessions | 6/6 (was 4/6) |
| Prompts answered | 8 — 6 resume-from-summary, 2 folder-trust |
Judgements made, and why
- "Resume from summary" over "full session" on every large session (544.1k, 418.2k, 170.1k tokens). A full resume consumes a large slice of the weekly limit for no gain.
- Answered the
/Users/richhtrust prompts "Yes" only after Rich said "all should resume accordingly". My first instinct was to close those panes rather than grant home-directory trust; I changed course on his explicit instruction, not on my own. - Did NOT fire TCC request prompts unattended. A
timed-out prompt writes
flags=1, which greys the checkbox permanently. That is the one irreversible action here, so it stays a documented one-liner for a human at each keyboard. - Did NOT edit PATH or overwrite Homebrew's link on the two broken hosts. Moving permanent global state to work around a symptom nobody understands is a bad trade; the fix is contained to one script instead.
- Did NOT tear down mesh links into rdmsm4x on an untested exhaustion hypothesis — Rich asked for every host to attach to every other host, and MaxStartups counts unauthenticated connections, which established links are not.
Two hosts repaired (HERDR-20260905-06)
claude hung before executing on rdmbair13m5/jdmbair13m5
— stuck at _dyld_start, footprint frozen at 32KB. Narrowed
to the exact absolute path: byte-identical copies run
in 1s as …/2.1.261/zz-probe (same directory) or
/opt/homebrew/zzdir/claude (same filename), while
…/2.1.261/claude never returns. Five hypotheses disproved
by measurement — contention (a lone launch with zero other claude
processes still timed out), quarantine (xattr -r -d to
zero), inode (replaced with a copy of itself), ~/.claude
state (clean HOME), and code signature
(codesign -v 0s and identical on working hosts).
brew reinstall did not fix it, and on jdmbair13m5 it left
the binary mode 644 (rc=126). Root cause still unexplained. Workaround:
working copy at ~/bin/claude;
fleet_herdr_resume.zsh probes for a binary that actually
returns a version. Both hosts now run 4 resumed sessions each.
Two fleet faults found along the way
- HERDR-20260905-08 — stale duplicate Tailscale
nodes.
rdmbair13m5/rdmbair15m5resolve to devices last seen 15 days ago; the live Macs re-registered as-1. Cost 2 of 30 mesh links. ssh aliases added on jdmbair13m5; the stale tailnet nodes should be deleted at source. - HERDR-20260905-10 — rdmsm4x sshd refuses every
connection (host UP: ping 2/2, TCP :22 accepts then resets,
Tyrell manifest fresher than any other host; census claude 30 / codex 10
/ gemini 6). Canonical pushes are blocked, so this changelog,
ISSUES.mdandSESSION-STATE.mdare LOCAL ONLY until it recovers.
An ENFILE alarm on this host was raised against the herdr build-out
and disproved: two runaway gopls held
1,037,566 and 1,037,524 descriptors — 99% of the kernel table — while
every claude/herdr process held ~350. Traced by a peer session to agy
mirror clients, filed as their ISS-006. herdr accounted for roughly
0.07%.
Corrections to earlier claims in this session
- "Killing the 31 orphaned daemons made it worse (112s → 146s)" is
void — both figures were
timeoutvalues on a command that never completes. Two timeouts are not durations. - A mesh count of "2/5 on rdmbair15m5" was a measurement
artifact: it counted only
shell@labels while three peers were still attached under phase-1claude@<peer>labels. Recounted correctly as 5/5. I nearly "healed" something that was not broken. ps | awk /claude --version/killed the ssh session running it, because that session's own argv contains the string.pkill -x(exact comm match) is the safe form.rc=$?after a pipe reports the LAST stage's status; two of my probes readhead's exit code.
New/changed
scripts (canonical ~/dev/fleet/herdr/, push pending)
fleet_herdr_mesh.zsh— builds the mesh + per-host resume, with a health guard that skips resume whereclaudewill not start.fleet_herdr_resume.zsh— resumes local sessions by explicit id; never resumes or closes its own workspace; skips sessions with no recoverablecwd(they would trigger a$HOMEtrust prompt); probes for a working claude binary.fleet_herdr_unblock.zsh— allow-list prompt answerer. Answers only the two known prompts; anything unrecognised is reported and left alone. It must never become "press Enter on whatever is blocking".fleet_herdr_attach.zsh— now idempotent; safe to re-run as a heal.
No secrets written to any file, note or log.