rdmsm4x-changelog-20260905-1139-prd-durability-and-load-root-cause
Session: claude@rdmsm4x/879ddeaa · 2026-09-05 11:30:50 - 11:39:25 EDT Trigger: Rich, "resume all work" Summary: Closed the verified-and-uncommitted window on 111 generated project PRDs (0 now at risk), cleared a stale git index.lock that had silently blocked one repo for 2d19h, and identified the root cause of the fleet's DEGRADED delivery grade as launchd job starvation under load 126 on rdmsm4x.
Scope
Host: rdmsm4x only. No remote host was modified. No process was killed. No config changed.
What changed
1. 111 generated project PRDs reconciled to zero-at-risk (56 commits)
FEAT-20260904-05 M2 generated 113 PRDs on 2026-09-04
03:00 and the ticket was marked resolved - but the
artifacts were never committed. That is the documented
verified-and-uncommitted state that gets silently reverted in a shared
tree.
Final accounting (counts, not "green"): | state | n | |---|---| |
total PRD files on disk | 111 | | tracked at exact path | 88 | | tracked
under case-variant Docs/ | 11 | | deliberately gitignored |
5 | | still untracked / at risk | 0
|
- 56 repos committed, path-scoped
(
git add -- docs/project_prd.html).~/dev/fleetalone had 213 dirty files from other agents;git add -Athere would have swept them into this commit. - 11 were already committed by a peer agent as
Docs/project_prd.html(capital D) during this session - concurrent work on the same files, not a conflict. - 5 are deliberately gitignored (
lib/arista,lib/xcode-config,lib/terminal_color_profiles,lib/templatesare ignored dirs;fleet/macosignores the file by name). NOT force-added - overriding a deliberate exclusion is not mine to do.
2. Stale
.git/index.lock cleared - sites/dev.dataroo.net
~/dev/sites/dev.dataroo.net/.git/index.lock, 0 bytes,
created 2026-09-02 16:49:22, i.e. 2d 19h old. Verified
stale before touching: lsof on the lock returned rc=1 with
no output, and no git process referenced that repo. It had been silently
failing every write to that repo. Moved, not deleted,
to the session scratchpad:
.../scratchpad/dev.dataroo.net-index.lock.stale-20260902-1649
Repo confirmed writable afterwards; the blocked PRD committed as
7ea245f.
Findings filed (routed, not absorbed)
- ISSUE-20260905-08 (incident, high) - rdmsm4x load 126/16 cores starves launchd StartInterval jobs; root cause of the DEGRADED fleet delivery grade.
- ISSUE-20260905-09 (bug, low) - 11 app repos track
Docs/project_prd.htmlwith capital D; invisible on APFS, breaks on case-sensitive filesystems. No data at risk.
Root cause of the DEGRADED grade (ISSUE-20260905-08)
Answering high-priority bus message
20260905-102819-34BD39EA (stability review, fleet 73%). The
per-job ratios are one fault with N symptoms, not N faults:
load 126 on 16 cores -> launchd defers StartInterval jobs (drops missed runs, does not
queue them) -> every interval job under-delivers at once -> dashboards/heartbeats stale
-> review grades fleet DEGRADED
Measured, not inferred:
launchctl print gui/501/com.eastcoastscience.devsite:runs = 23,run interval = 1800s, uptime 29h => expected ~58, delivered 23 = 40%.last exit code = 0,state = not running. A healthy job and a starved job report identically; the run COUNT is the only tell.- Load is NOT compute: sum(%CPU) = 399.5 of 1600 available (25% utilised), 1321 processes, 58 runnable. It is process-count and FSEvents contention.
- Anomalies: ~28x
git fsmonitor--daemon(one per repo, >30h each, ~80 repos in one tree); 80xSetStoreUpdateService(ppid=1);LoginItems.appexpinned 99% for 14h11m; 26 zombies;mediaanalysisd343%,contactsd79%; 23x mem0-mcp / 11x firecrawl-mcp / 9x tyrell-mcp. - rdmpw3265m/rdmpw3275m at 64% are the same mechanism under Rich's authorised build load.
Corrected a would-be false alarm: config-sync reports
diverged=0 on every run since 2026-09-04
17:49. ~/.agent-coordination/config-sync.escalated is a
stale dedupe ledger, not live state. Intermittent "offline" hosts are
SSH timeouts under this same load.
Verification evidence
- Zero-at-risk recount run AFTER all commits, with a case-insensitive tracking lookup (the case-sensitive version produced a false "17 at risk" reading).
- Lock staleness proven by
lsofrc=1 + absent process, before the move. - devsite starvation read from
launchctl printrun counts, not from exit codes.
How to undo
- PRD commits:
git -C <repo> revert <sha>per repo, orgit rm --cached docs/project_prd.html. All 56 commits share the subjectdocs: add generated project PRD. - index.lock: restore from the scratchpad path above (restoring it re-blocks that repo).
Outstanding owner actions (Rich)
- git fsmonitor - are ~28
fsmonitor--daemonprocesses over ~80 repos wanted? Most likely primary load driver. Reversible:git config --unset core.fsmonitorper repo. - macOS service leak -
LoginItems.appexat 99% for 14h and 80xSetStoreUpdateServicelook genuinely stuck; targeted restart is a candidate but they are long-lived processes and ACTIVE-DIRECTIVES 2026-09-04 21:10 forbids disrupting active user work. - Which in-progress work to drive next - 14 in-progress, 9 blocked, 3 needs-decision tickets.
Budget
Anthropic proceed (5h 2%, 7d 22%). OpenAI
hold at 100%, resets 2026-09-07 02:25 UTC - no
codex jobs until then. Aggregate advice was hold
on the OpenAI bucket alone; all work this session was done inline with
no subagents spawned.
Notes
No secrets written. Apple Notes entry PENDING - this
session is launchctl managername = Background,
which cannot send AppleEvents to the Notes GUI. File copy is the
archive; the Notes entry must be added from an Aqua session.
CORRECTION — appended 2026-09-05 17:29 EDT (same session)
rdmsm4x rebooted at 13:51:36 EDT, between the measurements above (11:30–11:40) and this append. Everything above was correct when taken but describes a machine instance that no longer exists. I did not detect the reboot until re-measuring at 17:27; the correction was sent to the session collecting owner-gated items before Rich could act on the withdrawn item.
Withdrawn
Outstanding owner action #2 (restart stuck macOS services) is
closed, not needed. LoginItems.appex
pinned at 99% for 14h11m is gone — the reboot cleared it.
Strengthened
Owner action #1 (git fsmonitor) is now better evidenced, because the accumulation regenerates:
| pre-reboot (30h uptime) | post-reboot (3h35m uptime) | |
|---|---|---|
git fsmonitor--daemon |
~28 | 31 |
SetStoreUpdateService |
80 | 71 |
| total processes | 1321 | 1387 |
| zombies | 26 | 20 |
| sum %CPU of 1600 | 399.5 (25%) | 816.2 (51%) |
| load1 | 126.49 | 23.38 |
31 fsmonitor daemons and 71 SetStoreUpdateService rebuilt in 3h35m versus 28/80 in 30h. These are not one-off leaks a reboot resolves. The reboot bought time, not a fix.
Acute starvation is relieved for now (ratio 7.91 → 1.46). The launchd delivery ratios must be re-graded against the new 3h35m window, not the old 29h one — grading post-reboot jobs against pre-reboot uptime yields a false failing grade, the exact error the stability review warns about. I have not re-graded them.
Committed work verified to have survived the reboot
Recounted at 17:28, against the 11:39 baseline recorded above:
| 11:39 | 17:28 | |
|---|---|---|
| total PRD files | 111 | 112 (one newly generated since) |
| tracked at exact path | 88 | 89 |
tracked under Docs/ |
11 | 11 |
| deliberately gitignored | 5 | 5 |
| at risk | 0 | 0 |
sites/dev.dataroo.net/.git/index.lock did not reappear;
commit 7ea245f still HEAD-reachable. Recording counts
rather than "green" is what makes this delta checkable at all.
New, uninvestigated — flagged, not asserted
1Password at 393% CPU over 2h49m and
spotlightknowledged.updater at 71%, both post-reboot.
Possible next load source. No investigation performed; not raised as
findings.
Notes entry RESOLVED — appended 2026-09-05 17:31 EDT
The "Apple Notes entry PENDING" note above is superseded — it was published.
Published by peer session dev-f6
(claude@rdmsm4x/88af4d) at 17:29:14, rc=0, into Notes / llmlog
as
rdmsm4x - changelog - prd-durability-and-load-root-cause - claude - unspecified - general - 20260905-1139.
dev-f6 confirmed it with an AppleScript
whose name contains "prd-durability" query returning 1 hit,
not merely the script's exit code.
Provenance, stated honestly: this is dev-f6's
verification, not mine. This session is
launchctl managername = Background and cannot
send AppleEvents, so I could not independently confirm the note exists.
Recorded as a peer-verified claim.
Routing lesson (mine). Being Background
blocked me from publishing, and I escalated that to Rich as an
owner-gated item. It was not owner-gated at all — dev-f6 was
Aqua (launched from a Terminal at the console) and simply
did it. "Agent sessions are Background" is not universally true; it
depends how the session was started. The correct move when blocked by
session TYPE is to check for a peer with the right type and route to it,
and only escalate to Rich if none exists.