Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260905-1139-prd-durability-and-load-root-cause

rdmsm4x-changelog-20260905-1139-prd-durability-and-load-root-cause

Session: claude@rdmsm4x/879ddeaa · 2026-09-05 11:30:50 - 11:39:25 EDT Trigger: Rich, "resume all work" Summary: Closed the verified-and-uncommitted window on 111 generated project PRDs (0 now at risk), cleared a stale git index.lock that had silently blocked one repo for 2d19h, and identified the root cause of the fleet's DEGRADED delivery grade as launchd job starvation under load 126 on rdmsm4x.

Scope

Host: rdmsm4x only. No remote host was modified. No process was killed. No config changed.

What changed

1. 111 generated project PRDs reconciled to zero-at-risk (56 commits)

FEAT-20260904-05 M2 generated 113 PRDs on 2026-09-04 03:00 and the ticket was marked resolved - but the artifacts were never committed. That is the documented verified-and-uncommitted state that gets silently reverted in a shared tree.

Final accounting (counts, not "green"): | state | n | |---|---| | total PRD files on disk | 111 | | tracked at exact path | 88 | | tracked under case-variant Docs/ | 11 | | deliberately gitignored | 5 | | still untracked / at risk | 0 |

2. Stale .git/index.lock cleared - sites/dev.dataroo.net

~/dev/sites/dev.dataroo.net/.git/index.lock, 0 bytes, created 2026-09-02 16:49:22, i.e. 2d 19h old. Verified stale before touching: lsof on the lock returned rc=1 with no output, and no git process referenced that repo. It had been silently failing every write to that repo. Moved, not deleted, to the session scratchpad: .../scratchpad/dev.dataroo.net-index.lock.stale-20260902-1649 Repo confirmed writable afterwards; the blocked PRD committed as 7ea245f.

Findings filed (routed, not absorbed)

Root cause of the DEGRADED grade (ISSUE-20260905-08)

Answering high-priority bus message 20260905-102819-34BD39EA (stability review, fleet 73%). The per-job ratios are one fault with N symptoms, not N faults:

load 126 on 16 cores -> launchd defers StartInterval jobs (drops missed runs, does not
queue them) -> every interval job under-delivers at once -> dashboards/heartbeats stale
-> review grades fleet DEGRADED

Measured, not inferred:

Corrected a would-be false alarm: config-sync reports diverged=0 on every run since 2026-09-04 17:49. ~/.agent-coordination/config-sync.escalated is a stale dedupe ledger, not live state. Intermittent "offline" hosts are SSH timeouts under this same load.

Verification evidence

How to undo

Outstanding owner actions (Rich)

  1. git fsmonitor - are ~28 fsmonitor--daemon processes over ~80 repos wanted? Most likely primary load driver. Reversible: git config --unset core.fsmonitor per repo.
  2. macOS service leak - LoginItems.appex at 99% for 14h and 80x SetStoreUpdateService look genuinely stuck; targeted restart is a candidate but they are long-lived processes and ACTIVE-DIRECTIVES 2026-09-04 21:10 forbids disrupting active user work.
  3. Which in-progress work to drive next - 14 in-progress, 9 blocked, 3 needs-decision tickets.

Budget

Anthropic proceed (5h 2%, 7d 22%). OpenAI hold at 100%, resets 2026-09-07 02:25 UTC - no codex jobs until then. Aggregate advice was hold on the OpenAI bucket alone; all work this session was done inline with no subagents spawned.

Notes

No secrets written. Apple Notes entry PENDING - this session is launchctl managername = Background, which cannot send AppleEvents to the Notes GUI. File copy is the archive; the Notes entry must be added from an Aqua session.


CORRECTION — appended 2026-09-05 17:29 EDT (same session)

rdmsm4x rebooted at 13:51:36 EDT, between the measurements above (11:30–11:40) and this append. Everything above was correct when taken but describes a machine instance that no longer exists. I did not detect the reboot until re-measuring at 17:27; the correction was sent to the session collecting owner-gated items before Rich could act on the withdrawn item.

Withdrawn

Outstanding owner action #2 (restart stuck macOS services) is closed, not needed. LoginItems.appex pinned at 99% for 14h11m is gone — the reboot cleared it.

Strengthened

Owner action #1 (git fsmonitor) is now better evidenced, because the accumulation regenerates:

pre-reboot (30h uptime) post-reboot (3h35m uptime)
git fsmonitor--daemon ~28 31
SetStoreUpdateService 80 71
total processes 1321 1387
zombies 26 20
sum %CPU of 1600 399.5 (25%) 816.2 (51%)
load1 126.49 23.38

31 fsmonitor daemons and 71 SetStoreUpdateService rebuilt in 3h35m versus 28/80 in 30h. These are not one-off leaks a reboot resolves. The reboot bought time, not a fix.

Acute starvation is relieved for now (ratio 7.91 → 1.46). The launchd delivery ratios must be re-graded against the new 3h35m window, not the old 29h one — grading post-reboot jobs against pre-reboot uptime yields a false failing grade, the exact error the stability review warns about. I have not re-graded them.

Committed work verified to have survived the reboot

Recounted at 17:28, against the 11:39 baseline recorded above:

11:39 17:28
total PRD files 111 112 (one newly generated since)
tracked at exact path 88 89
tracked under Docs/ 11 11
deliberately gitignored 5 5
at risk 0 0

sites/dev.dataroo.net/.git/index.lock did not reappear; commit 7ea245f still HEAD-reachable. Recording counts rather than "green" is what makes this delta checkable at all.

New, uninvestigated — flagged, not asserted

1Password at 393% CPU over 2h49m and spotlightknowledged.updater at 71%, both post-reboot. Possible next load source. No investigation performed; not raised as findings.


Notes entry RESOLVED — appended 2026-09-05 17:31 EDT

The "Apple Notes entry PENDING" note above is superseded — it was published.

Published by peer session dev-f6 (claude@rdmsm4x/88af4d) at 17:29:14, rc=0, into Notes / llmlog as rdmsm4x - changelog - prd-durability-and-load-root-cause - claude - unspecified - general - 20260905-1139. dev-f6 confirmed it with an AppleScript whose name contains "prd-durability" query returning 1 hit, not merely the script's exit code.

Provenance, stated honestly: this is dev-f6's verification, not mine. This session is launchctl managername = Background and cannot send AppleEvents, so I could not independently confirm the note exists. Recorded as a peer-verified claim.

Routing lesson (mine). Being Background blocked me from publishing, and I escalated that to Rich as an owner-gated item. It was not owner-gated at all — dev-f6 was Aqua (launched from a Terminal at the console) and simply did it. "Agent sessions are Background" is not universally true; it depends how the session was started. The correct move when blocked by session TYPE is to check for a peer with the right type and route to it, and only escalate to Rich if none exists.