herdr mesh: --hosts flag defect fixed, and a hub-saturation reading corrected
Session: claude@rdmbair15m5 (ca8a3ca6) · 2026-09-25 10:11:32 → 10:21:03 EDT · post-reboot resume, prompted by grok@rdmsm4x/grok0924
What I resumed into
Re-read my check-in trail and bus inbox. Two things shaped the work:
- claude@rdmpw3275m/d5f5ef68 had already restored the full
mesh (TASK-20260924-21, 2026-09-24 23:28) — so my recorded
"re-verify the mesh" step was mostly done. Their check-in's resume note
and
do_not_touchlist were followed. - claude@rdmsm4x is priming Full Disk Access for
ECSToolHost (bus advisory
20260924-220411-1B79F5F3, relayed from Rich) and asked that nobody touch TCC/FDA or the herdr/ECSToolHost LaunchAgents until reported done. No completion is posted anywhere as of 2026-09-25 10:21:03 EDT, so all TCC verification was deliberately deferred — see HERDR-20260925-24.
Mesh state (measured, not assumed)
30/30 links across six hosts. jdmbair13m5 was at 4/5 on arrival, missing the link to the hub.
Probe note: the recorded check is
pgrep -f 'ssh .*-R .*herdr', but pgrep -f
matches the shell holding the command, so it counts itself. I fed the
probe over stdin with the pattern assembled at runtime.
First attempt returned empty on all five remotes because I used
ssh -n and a heredoc — -n redirects
stdin from /dev/null, so nothing reached the remote shell. That reads
exactly like "no mesh" when it is really "no probe".
HERDR-20260925-22
— --hosts emptied the peer universe (FIXED, commit
d12a46f)
ALL_HOSTS was both the operate-on list
and the peer universe, and --hosts overwrote it, so
--hosts "jdmbair13m5" computed
peers=(${ALL_HOSTS:#jdmbair13m5}) = EMPTY:
a single-host mesh repair attached nothing and still printed a pass.
Split into FLEET (immutable peer universe) and
ALL_HOSTS (hosts to operate on). Verified both
directions — single-host now lists all five peers; full fleet
unchanged at 6 hosts x 5 peers.
Same commit: the summary counted
grep -c 'pane running ssh' — attach
attempts, not live links — so a host whose link died
still read as complete. It now also counts live ssh -R
links and WARNs when live < peers. Commit is 1
file; ISSUES.md and the other lane's uncommitted
paths were deliberately left alone and snapshotted instead.
HERDR-20260925-23 — my hub-saturation reading was an lsof artifact
jdmbair13m5 -> rdmsm4x failed 3/3 with
kex_exchange_identification: Connection reset by peer. I
measured the hub with lsof -iTCP:22 -sTCP:ESTABLISHED and
read 265 connections, 84 from rdmpw3275m, and began
narrating saturation.
Wrong — the client side caught it. rdmpw3275m had
3 outbound ssh processes and 2 mem0 bridges. lsof emits
multiple ROWS per connection. Deduped by socket pair: 30
connections, 2 from rdmpw3275m. Configured on the hub:
MaxStartups 10:30:100, and the launchd job reading
not running / active count 0 (normal for socket
activation). So it is MaxStartups refusing pre-auth during a
burst, transient — once the hub was back to ~15 connections the
same link succeeded 2/2 with positive controls from two other hosts. The
memory entry framing this as a "launchd instance cap of 42" matched no
reading taken today and was corrected.
I did not kill anything. The three mem0 bridges on this host all had live parents (Claude.app 1229, ChatGPT.app 1271) — none orphaned — so reaping them would have broken other lanes' running apps for no gain.
Credit where due
The dropped link self-healed while I was diagnosing:
another lane's commit
76f1817 fix(herdr): launcher auto-reconnects, so the fleet mesh survives sleep
regenerated attach-rdmsm4x.zsh on jdmbair13m5 at 10:15 and
the link came back. That was their fix working, not mine.
Records
fleet/herdr/ISSUES.md 27 headings (was 24), snapshotted
at
~/.agent-coordination/snapshots/rdmsm4x/fleet/20260925-102016
· memory hub-sshd-instance-cap-mem0-bridge corrected +
index reworded · commit d12a46f on canonical
~/dev/fleet main. No secrets written.