rdmsm4x-changelog-20260925-0015-grok-fixall-pass
rdmsm4x changelog 2026-09-24 — grok fix-all pass (grok@rdmsm4x/grok0924)
Operator: grok (Grok Bot executor), driven from rdmbair15m5 over ssh. Standing permission from Rich for reversible fixes. Hard limits observed: no reboot, no sudo/console, no key rotation, no irreversible deletion, no agent-session kills, RTTy 523 gate untouched.
1. Hub hang — ISSUE-20260924-33 (RESOLVED)
Contributors: Time Machine first full backup (5.3 TB) to smbfs NAS unaspro818a (backupd/mds_stores in U state, ~190 MB/s), lsof stat'ing the smbfs mount, CleanMyMac_5_MAS 454% CPU for 9 h (exited by itself), wedged launchd monitors.
- dev/scripts d02362a — resume_agent_sessions.zsh v1.0.1:
timeout-bounded lsof/ps,
lsof -w, no zero-arg lsof. Snapshot 6.5 h wedge -> 5.6 s. Deployed on all 6 hosts (backups*.bak-grok-*). - fleet 91437ef — tooling/host-scripts/probe_host_state.sh v2.1: one lsof pass. 30+ min -> 7.6 s. Deployed hub ~/scripts + rdmbair15m5 only (other spokes run md5 a0cbd34e, left alone).
- dev/scripts 6e48fbd —
grok0924-guard(zsystem flock no-overlap lock in ~/.agent-coordination/locks/.lock; ps/pgrep/lsof under timeout 30/30/60) in fleet_audit.zsh, fleet_stability_review.zsh; same block in ~/.agent-coordination/agent_status.zsh + agent_heal.zsh on hub, rdmbair15m5, rdmbair13m5, rdmpw3265m, rdmpw3275m (not jdmbair13m5: different agent_status, no agent_heal). Tested: free run rc0 + heartbeat; held-lock run skips rc0. - Stopped wedged monitor runs only (snapshot pids 23077, 28218; fleetstate probe 22852/38245). Nothing paused/unloaded.
- Evidence: ~/dev/_handoff/grok-fleetops-20260924/evidence/ (shutdown_stall_225612.txt, log-19-2250.txt).
- Undo: restore
*.bak-grok-*copies /git revert d02362a 6e48fbd(scripts),git revert 91437ef(fleet).
2. fileproviderd — ISSUE-20260905-08 (commented, released, left in-progress for overnight recheck)
- issues repo 270da51 (merge of aaac8ac): bin/ticket local HTML -> ~/Library/Application Support/ticket/html (override TICKET_LOCAL_HTML_DIR). Pushed fleet + backup.
- Tests: 14 test scripts all rc0 (e2e 89/89, fts5 77/77, jsonrpc 31/31, sync-trailers 76, status-reopen 6/6).
- Migration: 1,219 files copied; ~/Desktop/issues moved to ~/Library/Application Support/ticket/desktop-issues-moved-20260924-2251; ~/Desktop/Tickets dashboard.webloc added.
- Post-reboot: fileproviderd ~0% CPU; fsmonitor daemons 29 -> 5.
- Undo:
git revert -m1 270da51; move the desktop-issues-moved dir back to ~/Desktop/issues.
3. Bus expire — agent_msg.zsh v1.6.0
- maintenance repo ba3fb48, merged 5021d4b, pushed fleet + backup. Tests 67/67 (baseline 52/52; 15 new §12 checks — commit msg says +16, typo).
expire/unexpiretombstones in .state (reversible, not an ack); inbox hides expired,inbox --expiredshows.- Live on all 6 hosts (md5 05cb1be3…), v1.5.2 backed up alongside.
- Expired 675: 642 superseded fleet-audit/stability-review snapshots (newest per family kept) + 33 undeliverable (dead identities). Lists: ~/dev/_handoff/grok-bus-triage-20260924/expired-*-20260924-2310.txt. claude@rdmsm4x NEW-high 676 -> 37.
- Undo:
agent_msg.zsh unexpire --file <list>.
4. ecsmem0d priority — TASK-20260924-16 (commented, not claimed)
- Hub plist: ProcessType Standard, Nice/LowPriorityIO removed. Backup ~/Library/LaunchAgents/.bak/com.eastcoastscience.ecsmem0d.plist.bak-grok-20260924-2252. /health 200 in 0.095 s; survived reboot.
- SelfBootstrap Swift writer will revert it — Tyrell lane.
5. mem0 bridge
- ~/bin/mem0-mcp v2 on all 5 spokes: own ssh connection (ControlMaster=no, ControlPath=none, ServerAlive 30x4). Backups ~/bin/mem0-mcp.bak-grok-20260924-225x. initialize OK on each.
6. firecrawl MCP — ISSUE-20260924-31
- Cause: spokes had no working npx (config is
npx -y firecrawl-mcp). - rdmpw3265m + rdmpw3275m: brew node 26.10.0 built from source
(Intel);
claude mcp list= firecrawl, mem0-shared, replicantdb, ticket-mcp, tyrell all Connected on both. - Pending (Airs not rebooted by 00:12, brew deferred): jdmbair13m5
brew reinstall node(dangling /opt/homebrew/bin/npx); rdmbair13m5brew install node. fnm node v24.21.0 fallback on all spokes. Ticket released with notes.
7. Tyrell Tailscale fixture — ISSUE-20260924-32
- Branch grok/tailscale-fixture-ISSUE-20260924-32 f96fa48 (fixture-*.tyrell-test.invalid, 192.0.2.x). Worktree ~/dev/_worktrees/tyrell-grok-fixture-20260924.
- rdmbair13m5 full
swift teston branch: Executed 184 tests, 0 failures; testTailscaleJSONParsing passed. - NOT merged: --no-ff merge onto fleet/main a513bc5 (local c666c7a, branch grok/merge-fixture-ISSUE-20260924-32 on rdmbair13m5) fails to link TyrellCoreTests. Cause on main: OfflineChatQueueTests.swift (2172451, FEAT-20260831-17) imports TyrellAppSupport/TyrellRemoteSupport, not in TyrellCoreTests deps. The push to main also got an approval block. Left for the Tyrell lead. Ticket released.
- Undo: delete branches grok/tailscale-fixture-ISSUE-20260924-32 /
grok/merge-fixture-ISSUE-20260924-32;
git worktree remove ~/dev/_worktrees/tyrell-grok-fixture-20260924(on rdmbair13m5 and hub).
8. Pushes
- maintenance 9bfda0b..5021d4b -> fleet + backup. dev/scripts a0bac97..01cf850 -> fleet (backup already current; includes two non-grok commits fe8c7af, 01cf850 already on backup). fleet repo 0/0. issues 270da51 on both.
9. Post-reboot verification
- Hub (rdmsm4x, 27.2 26B5091g, unchanged): all 43 eastcoastscience jobs loaded; ecsmem0d :8767/health 200; tyrelld up; ticket-mcp 3.2.0 OK; overnight jobs (fleetupdate, devbackup, certsync, nightlyaudit, fleetaudit, stabilityreview) loaded rc0. Load 547 at boot -> ~43 at 23:47. fileproviderd 0% (30 s CPU in 49 min). TM backup copying (0.4%).
- rdmpw3265m / rdmpw3275m: macOS 26.7.1 (25G313), same as FLEET.md, so no FLEET.md edit. tyrelld up, ecsmem0d 200, fleet jobs rc0, bus v1.6.0.
- mem0 bridge initialize OK from rdmbair13m5, jdmbair13m5, rdmpw3265m, rdmpw3275m (00:00).
- Airs (rdmbair13m5, jdmbair13m5, rdmbair15m5) NOT rebooted as of 00:12. Their postboot check is still to do (script /tmp/grok0924/postboot.zsh on rdmbair15m5).
- Pre-existing launchd disables (not grok): stormguard, fleet-load-guard, fleet-realtime-guard, claude-quota-monitor, replicantDB.daemon.
Still broken / not mine
- devsite (publish_site) FAIL rc=124 every run since 14:07 incl. 23:20 post-boot: projects/site/ecs0 stages hit timeouts (gen_projects runs git per repo serially; stalls under IO contention from TM backup). Was 34-180 s/stage this morning.
- Hub TM full backup still running (load).
- RTTy 523 gate ended rc=1 at 22:52:21 (qa_release-523.rc=1), no v0.3.848 tag — lead's call.
- TASK-16 writer revert; jdmbair13m5 guard not deployed; spoke probe_host_state versions differ.
Rich-only
- Apple Notes entry (Rule 26) from a GUI session:
zsh ~/scripts/notes_changelog.zsh <this file>. - Chrome flags reset (rdmpw3275m, jdmbair13m5); jdmbair13m5 lock delay 300 s -> Immediately; GPM passkey/PIN settings.
- TM: let run /
tmutil stopbackup/ schedule off-hours. - Optional: sshd MaxSessions (sudo).