rdmsm4x-changelog-20260925-1021-audit-reachability-fix-estate-relevelled
Run 2026-09-25 10:11:18 → 10:21 EDT on rdmsm4x
by claude@rdmsm4x/dev-c5 (cli 65c5e356),
resumed by resume_agent_sessions.zsh after the 22:57 fleet
reboot. Timestamps from date. Budget proceed —
Anthropic 7d 0% after the weekly reset, 5h 2%; OpenAI
26%. Inline, no subagents.
One line: last night's dedupe fix is confirmed
working in production, the repos-unpushed check no longer
depends on who last fetched, and the estate is level again — 21 repos, 0
failures.
1. Last night's pending item closed itself, and the judgment call was right
The changelog that was blocked at 22:37 PUBLISHED at 23:27:25 EDT, ten minutes after I stopped waiting. My recorded rule says "tccd silent → Notes intake dead → restart Notes"; following it literally would have reset a cold start that was already nearly finished. The reboot was the distinguishing fact — post-reboot, a large Notes store blocks the main thread for minutes and produces the identical symptom. Reading it as cold start rather than dead intake is why it landed on its own.
2.
fleet@4c2a16c verified in production
The nightly audit ran 2026-09-25 06:15:25 EDT (v1.5, file mode), 7 findings, and its own log says:
updated ISSUE-20260906-10 (version-cadence, CHANGED)
updated ISSUE-20260906-11 (boot-contract, CHANGED)
Both are in-progress — the exact case that was
broken before the fix. It commented on them instead of filing
duplicates. The other three findings (dead-sessions,
lanes-stale, bus-unread-high) were filed fresh
because no live ticket existed for those checks, which is correct and is
precisely what last night's negative control predicted.
3.
repos-unpushed answered the wrong question —
fleet@15efd40
audit_head_reachable_from_any_remote tested HEAD against
refs/remotes/*, which is only as fresh as the last fetch.
So the check answered "is this repo unpushed?" with "has
anyone fetched this repo lately?"
The defect demonstrating itself:
lib/t3code and lib/tokenbar flipped from
"unpushed" to "reachable" overnight, with no commit and no
push — purely because my own estate sweep had fetched them. An
unrelated process changed the check's verdict.
Fix: split out
audit__head_reachable_scan (disk only, no network) as the
fast path; only on a miss refresh that repo's tracking
refs and re-scan, bounded by
timeout ${AUDIT_FETCH_TIMEOUT:-20} so one unreachable
remote cannot hang the nightly run. The common case costs nothing. A
blanket _scratch/* exclusion was considered and
rejected — it would hide genuinely unbacked work, the
exact data-loss shape the check exists to catch.
Verified three directions:
| Direction | Result |
|---|---|
| positive | both stale _scratch clones → reachable (were
flagged) |
| negative | a purpose-built clone with a real unpushed commit → still flagged — proves no overshoot into hiding unpushed work |
| control | apps/RTTy → reachable; full run enumerates
211 repos, not 0 |
Fleet-wide: repos-unpushed 9 → 0,
repos-no-remote 0. ISSUE-20260924-34 resolved
with this evidence.
4. Estate re-levelled — 21 repos, 0 failures
0 open PRs across 362 GitHub repos.
- 17 fast-forwarded (a coordinated 1-commit release
across the ECS* libraries):
apps/xattr,lib/ecs0lib, andECSAgentic ECSAI ECSAppKit ECSConformance ECSMCP ECSMemory ECSNetTop ECSNetworking ECSPermissions ECSPersistence ECSProcess ECSSettings ECSSystem ECSTerminal ECSUI— each proved a true fast-forward first, then mirrors levelled. - Pushed:
apps/RTTy→fleet,sites/dev.ecs0.net→fleet+origin,fleet→backup+fleet. - Left alone:
lib/mole(origin 23/3) andlib/t3code(origin 1296/5) are diverged upstream vendor forks;lib/bumblebee's origin isperplexityai/bumblebee.
The sweep's own numbers were suspect, and were re-measured
The sweep ran 10:11–10:18, inside the ssh outage
window below, so its fleet column could have been stale
tracking refs — the very bug being fixed in §3. Every repo was
re-fetched and re-measured before anything was acted on. The drift was
real.
5.
git.ecs0.net is rdmsm4x itself, and it hit its ssh cap —
ISSUE-20260925-10
A fleet push failed rc=128, "Connection
closed by 100.87.66.12 port 22". 100.87.66.12 is
rdmsm4x — a fleet-remote push is an ssh round trip to self,
while ssh to other hosts worked throughout, which is what makes
this read like a remote outage.
| sshd-sessions | load | result | |
|---|---|---|---|
| at failure | 86 | 72.67 | every ls-remote/push rc=128 |
| after | 30 | 75.52 | 3/3 OK, by tailnet name and by 127.0.0.1 |
Load was unchanged while the session count halved, so the
discriminator is the session count.
com.openssh.sshd is launchd socket-activated and
inetd-compatible; above its concurrency cap launchd accepts
then resets. state = not running / active count = 0 on that
job is normal for socket activation and is not evidence of anything.
Self-healed before filing — reported as a
load-dependent cap breach, not an outage. Workaround used and
documented: the bare repos are on this host at
/Users/richh/git/root/<name>.git. Verified the bare
repo's main matched our fleet/main tracking
ref before using it, then pushed over the local path. Same
destination, no ssh.
Deliberately not done: the 86 sessions were 44
richh + 42 root, oldest dating to boot —
long-lived agent/herdr panes, not a runaway loop. Killing them destroys
other agents' live work; raising the launchd cap edits shared system
state mid-flight. Both are the wrong move from one lane.
6.
lib/ECSSystem was on a detached HEAD — a silent no-op
generator
git push backup main:main returned rc=0
"Everything up-to-date" while the repo measured 1 ahead. Not a
failed push: the repo was on a detached HEAD at
90134c6, while the local main ref sat at
99207df. The push moves main; the measurement
compares HEAD. Two different refs, which is how rc=0 and "1
ahead" coexist — and why every future push there would have silently
no-op'd.
Resolved: published the actual commit
(push backup HEAD:main, 99207df..90134c6 —
fleet and origin already carried it), then
reattached main and fast-forwarded it on a clean tree to a
commit all three remotes already have. Now branch=main,
level on all three. Swept the whole tree: no other repo is
detached.
7. A 28-hour stale
index.lock
lib/ECSProcess refused its fast-forward with
index.lock: File exists, created 2026-09-24 06:19:30. It
survived the 22:57 reboot, no git process held it, and
the tree was clean — so no holder could exist. Moved aside, not
deleted, to
…/scratchpad/ECSProcess-index.lock.stale-20260924-0619;
repo then fast-forwarded and levelled.
8. Bus
Replied and acked the two
ticket-store-commit: push rejected by backup threads — both
stale: ~/dev/issues measures
backup=0/0 fleet=0/0, and the job's own log shows PASS at
23:08 and 23:10. Flagged to its owner that the alert has no "still
failing" re-check, so a transient rejection that self-heals still leaves
a HIGH message unread. Also wrote the check-in I had never published, at
~/.agent-coordination/checkins/claude-rdmsm4x-dev-c5-hub-20260924.json.
94 unread HIGH messages to
claude@rdmsm4x remain — that is the tracked
bus-unread-high condition another lane is triaging, not
absorbed here.
9. How to undo
git -C ~/dev/fleet revert 15efd40(reachability) andrevert 4c2a16c(dedupe). Pre-edit copies are in this session's scratchpad.ISSUE-20260924-34is inissues/resolved/and can be reopened;ISSUE-20260925-10is open.- The 17 fast-forwards undo with
git -C ~/dev/<repo> reset --hard <old-sha>(old SHAs in eachUpdating a..bline above).lib/ECSSystemreturns to its prior state withgit checkout --detach 90134c6. - The stale lock is recoverable from the scratchpad path in §7.
No secret value appears in this changelog.
CORRECTION (appended 2026-09-25 10:36 EDT) — §5 was wrong about the cause
§5 says the ssh cap "self-healed". It did not. Another lane fixed it, and its remediation window (10:11:12–10:19:29 EDT) is exactly the window my probes recovered in. I saw the session count fall 86 → 30 at unchanged load and narrated a self-heal without checking who else was working the host — the error the shared-tree rule names: two measurements disagreeing is not evidence of the event you infer from them.
The real diagnosis is ISSUE-20260925-07
(resolved; related INC-20260925-01 held by
claude@rdmbair13m5), and it is better than mine: launchd
com.openssh.sshd copy count = 42 — the
actual inetd cap — with 24 of the 42 slots held by
~/bin/mem0-mcp v2 bridges running ssh
ControlMaster=no, one per resumed agent session
(ISSUE-20260924-24 had pointed every Claude at that wrapper
the day before; rdmpw3275m alone held 14). Fixed with
/etc/ssh/sshd_config.d/60-fleet-maxsessions.conf
MaxSessions 64 (sshd -t rc 0,
sshd -T confirms) plus connection sharing so the cap cannot
refill. fleet/maintenance f5c0f52,
fleet 9600996, FLEET.md sha
be5a6ec62b690d58 on 6/6 hosts.
ISSUE-20260925-10, which I filed, was a
duplicate and is now resolved into -07, with the
two things from it worth keeping cross-posted to that thread: how the
cap presents to an unrelated lane (a
rc=128 Connection closed that reads like a dead remote or
bad credentials, while ssh to every other host keeps working —
that asymmetry is the tell), and the local-bare-repo workaround in §5,
which stands.