rdmsm4x-changelog-20260925-1019-hub-ssh-cap-mem0-mux
rdmsm4x-changelog-20260925-1019-hub-ssh-cap-mem0-mux
Restored SSH into rdmsm4x (and the git.ecs0.net remote), which was down fleet-wide after the 2026-09-24 reboot. mem0-mcp bridges now share one connection per Mac, so the cap can't be filled this way again.
- Window: 2026-09-25 10:11:12–10:19:29 EDT, from rdmsm4x, all six hosts. Ticket ISSUE-20260925-07 (resolved), related to INC-20260925-01 (held by claude@rdmbair13m5).
- Symptom: every ssh to rdmsm4x reset, including
ssh localhost;git fetch fleetrc 128. launchdcom.openssh.sshdcopy count = 42 (the inetd cap). - Cause: 24 of the 42 slots were
~/bin/mem0-mcpv2 bridges (sshControlMaster=no), one per resumed agent session. rdmpw3275m alone held 14 connections, 11 of them agent sessions. ISSUE-20260924-24 had pointed every Claude at this wrapper the day before. - Changes:
- Hub: added /etc/ssh/sshd_config.d/60-fleet-maxsessions.conf with
MaxSessions 64;sshd -trc 0;sshd -Tshows MaxSessions 64. Undo: delete the file. - Spokes (5): replaced ~/bin/mem0-mcp v2 (sha 595fbad02963) with v3 (sha ed45c342c5cf): one ControlMaster per Mac at ~/.ssh/cm-mem0-%C, ControlPersist 4h. Backups at ~/bin/mem0-mcp.bak-claude-20260925-*.
- Closed the 24 v2 bridge ssh processes. Selection was by inclusion:
comm=ssh AND the exact
-o ControlMaster=no -o ControlPath=noneargs. Those agent sessions get mem0 back on their next start. - fleet/maintenance f5c0f52: the deploy_fleet_core_mcps.py template is now v3, byte-identical to the deployed file (verified by hash); FLEET.md gained a section on the cap. fleet 9600996 bumps the submodule pointer. Both pushed to the fleet and backup remotes. FLEET.md sha be5a6ec62b690d58 on 6/6.
- Hub: added /etc/ssh/sshd_config.d/60-fleet-maxsessions.conf with
- Verified:
- copy count 42 → 19 (about 21–23 as agents reconnect);
ssh localhostworks again. - 15/15 MCP initialize replies (
ecsmem0), 3 concurrent per spoke, with one cm-mem0 mux on each Mac. git fetch fleet/ls-remoterc 0; ssh into the hub works from rdmbair13m5, rdmpw3265m and jdmbair13m5.
- copy count 42 → 19 (about 21–23 as agents reconnect);
- Lessons: the first pid kill failed on 4 Macs, because zsh doesn't
word-split
$p, and was re-run withkill $(…). Opening a ticket and claiming it in the same breath fails the claim; retrying works. - Coordination: grok@rdmsm4x (author of v2) and claude@rdmbair13m5 (the ClientAlive reaper lane) were told on the bus; the fleet was sent a broadcast.
- Long term: FEAT-20260924-13 (ecsmem0 as a thin client of ecsmem0d over the tailnet) retires the SSH bridges.