rdmsm4x-changelog-20260905-1346-sshd-launchd-inetd-cap-diagnosis
2026-09-05 12:15–13:46 EDT · rdmsm4x · claude@rdmsm4x
Fleet-wide ssh outage on the hub diagnosed to a launchd per-job inetd instance cap, not the descriptor exhaustion two peer sessions suspected. No local state was changed — this session was measurement-only from start to finish. Nothing killed, no config edited, no process reaped.
Scope
Host measured: rdmsm4x only. Findings shared with two peer Claude sessions (Tyrell/rdmbair15m5 and the herdr fleet owner) over the cross-session bus.
What was wrong
ssh to rdmsm4x accepted TCP on 22 then reset before key
exchange, fleet-wide. Cause:
launchd[1] [system/com.openssh.sshd:] Could not create new instance of inetd service:
67: Too many processes # errno 67 = EPROCLIM
macOS socket-activates ssh — launchd holds
TCP *:22 (LISTEN) and forks a handler per connection, so
pgrep -x sshd returning 0 is normal, not a dead daemon.
launchd accepted then declined to fork, closing before any banner.
Every ordinary limit was healthy, which is what made
it hard: kern.num_files 14897/5000000 (0.3%), procs 1447/16000, richh
974/10666, memory 89% free, zero swap, sshd -T exit 0,
sshd_config unchanged since 2025-10-28. The evidence was in launchd's
log, not sshd's — a connection never forked writes nothing under
sshd.
Consumed by: every Claude/agy/codex/ChatGPT client
fleet-wide holds one long-lived
ssh richh@rdmsm4x ~/bin/mem0-mcp for its lifetime. 21 of 28
@notty sessions held mem0-mcp, oldest 22h54m. One hub ssh
slot per agent session does not scale past ~40.
Timeline (all EDT)
- 12:15:26.217 first rejection — true onset. 12:25 was when the fleet noticed.
- rdmbair15m5 hit ENFILE at 12:16, one minute after; concurrent independent failure of the same ~12:15 resume storm. Neither host caused the other.
- 12:34:41 peer's owner closed 13 agy mirror clients (not my action).
- 12:34:37.972 last rejection ever; zero afterwards.
- 12:35–12:36 ssh recovered; verified
ssh localhostandssh 100.87.66.12both exit 0.
Verification evidence
- 60-min baseline 12:39–13:39, 120 samples: copy_count min 32 max 38 mean 33.28 median 33; ssh OK 120/120; zero rejections in launchd log for the hour.
- Cap bound: 39–42. Served 38 without rejection, rejected at 42.
- Burst reconstruction from launchd per-instance UUID spawn/remove events (1226/1207 over 11:40–12:20): flat 2.5 min, then +8 in 46 seconds into the cap. The "76 rejections over ten minutes" is outage duration, not approach.
Corrections made in-session (all recorded)
- "Will re-break within the hour" — withdrawn; an hour of flat data showed a plateau, not drift. Burst-driven, not drift-driven.
- Watch threshold of 40 — unsafe, sits at/above the cap. Moved to 37.
- 30s sampling with a 2-sample debounce — could not have warned on a 46-second ramp. Replaced with 10s interval plus a rate rule (rise ≥3 in 30s), which needs no knowledge of the cap and would have fired ~46s ahead. A level rule cannot warn on a step starting below it.
Instrument traps hit (each returned a confident wrong number, none errored)
pgrep -fc— macOS pgrep has no-c; returned 0 while 43 processes ran. Use| wc -l.lsofoutput read throughhead -60— showed only launchd rows, implying handlers had lost their sockets. Both hold them (launchd keeps a copy; handler gets fds 7u/8u).lsof -iTCP:22matches both directions; raw counts conflate inbound with this host's outbound ssh. Split before counting.- (from peers)
netstat -anon macOS 27 emits no TCP rows while exiting 0;pgrep -x sshd-sessioncan never match because comm carries a suffix.
Files written (all additive, no overwrites)
~/dev/todo/items/20260905-1232-rdmsm4x-sshd-launchd-inetd-cap.md— full record, four appends.~/.claude/projects/-Users-richh-dev/memory/launchd-inetd-cap-breaks-hub-ssh.md+ MEMORY.md line.- This changelog.
Outstanding owner actions for Rich
- ControlMaster multiplexing for rdmsm4x in each
spoke's ssh config — takes the fleet from one hub slot per agent session
to one per host. Six-host config change, not applied,
his call. Two traps if applied:
ControlPersistneeds a finite timeout (an unbounded master is itself a permanent slot), and the control socket must be per-%r@%h:%pon local disk, since a stale socket makes ssh silently fall back to a fresh connection while the config reads correct. Verify withssh -O check, not by reading the config. - Fan-out bracket measurement — peer is arranging one around the next fleet-wide resume, to decide whether 10s polling suffices or this should be event-driven off the launchd log.
- Structural: the mem0 MCP transport makes hub ssh availability a function of how many agents are alive fleet-wide, taking out the fleet's own recovery path.
How to undo
Nothing to undo — no state was changed. The only running artifact is a read-only monitoring loop (background task, 10s interval, 60 min, self-terminating). It can be left to expire.
PENDING — Apple Notes entry NOT written
launchctl managername = Background for
this session, so it cannot send AppleEvents to Notes and
notes_changelog.zsh refuses by design. This file is the
archive copy; the Notes entry in the rdmsm4x folder is still
owed and should be added from an Aqua session with:
zsh ~/scripts/notes_changelog.zsh <this file>
No secrets, credentials, tokens or signed URLs appear in this record.