jdmbair13m5-changelog-20260925-1016-hub-ssh-slot-exhaustion
jdmbair13m5 - incident - hub ssh slot exhaustion diagnosed, duplicate fix stood down - claude - post-reboot resume - fleet/hub-ssh - 2026-09-25 10:16 EDT
[2026-09-25 10:11:22 → 10:16 EDT · jdmbair13m5 · claude@jdmbair13m5/c59e20b4]
Resume
- Relayed by grok@rdmsm4x/grok0924 for Rich: fleet reboot done, weekly
limit reset. usage_status advice
proceed: anthropic 7d 0%, 5h 2%, vendor-reported. - Last check-in for this host (2026-09-23) is COMPLETED. No pre-reboot
check-in was written, so there was no recorded next step.
ticket_mine: 0 claims, 0 planned.~/devhere is scratch with no git repo.
Incident found at 10:11: rdmsm4x refused every new ssh connection
- Symptom:
kex_exchange_identification: Connection resetover LAN, .local and Tailscale. ticket-mcp: "claim authority rdmsm4x unreachable". - Cause:
/System/Library/LaunchDaemons/ssh.plisthasinetdCompatibility Instances = 42, and all 42 were held. 26 of them were mem0-mcp v2 bridges (ControlMaster=no, one hub connection per agent session): rdmpw3275m 13, jdmbair13m5 5, rdmbair15m5 2, rdmpw3265m 2, rdmbair13m5 1, plus 3 more over IPv6. All bridges were live. Resources were fine (89% memory free). - Owner: claude@rdmsm4x, ISSUE-20260925-07 / INC-20260925-01. It
shipped mem0-mcp v3 (dedicated mux per host) and the hub's
60-fleet-maxsessions.conf/50-fleet-clientalive.confwhile I was still diagnosing.
What I changed
- 10:15:00: wrote a duplicate hub
/etc/ssh/sshd_config.d/060-fleet-maxsessions.conf(MaxSessions 64). - 10:15: moved it to
rdmsm4x:~/dev/archive/sshd-dup-conf-20260925/and chowned it to richh. Afterwards:sshd -trc 0, effective MaxSessions 64. - Nothing else. No bridges killed, no launcher edited.
Verified
- Hub launchd sshd instances: 11 (were 42). Fresh ssh via bare
rdmsm4xandrdmsm4x.ts.dataroo.net: rc 0. ticket-mcp answers. - Bus: 20260925-101454-B018341A (diagnosis) and 20260925-101534-75486F0E (stand-down; its "10:16/10:17" times were estimates, actual 10:15).