rdmsm4x-changelog-20260926-0056-overnight-mem0-truth-cache-tyrell-build33
rdmsm4x changelog — overnight fleet health: mem0 truth, mem0 caching, Tyrell Build 33
One line: Tyrell's "Mem0 cannot connect" on spokes was a dashboard defect, not an outage. It is fixed and shipped in Build 33 to 5/6 hosts, spoke silo memory daemons are retired, and hub memory search now caches (repeat 85 ms → 0.26 ms, model resident 24 h).
- When: 2026-09-26 00:02:39 → 2026-09-26 00:56:27 EDT
- Host / agent: rdmsm4x · claude@rdmsm4x/3f92c388 (Opus 5.5 lead; 2 Sonnet workers for the repo and service sweeps)
- Ticket: EPIC-20260926-01 (children TASK-20260926-02/-03/-04); resolved TASK-20260924-23, ISSUE-20260924-30, TASK-20260924-14
Scope (hosts)
rdmsm4x, rdmbair15m5, jdmbair13m5, rdmpw3265m, rdmpw3275m. rdmbair13m5 was offline all night (Tailscale last seen ~4 h at 00:12).
Changes
- mem0 diagnosis: agent memory worked from every reachable spoke (MCP mem0_health through ~/bin/mem0-mcp → hub, 27 vectors). The Tyrell dashboard read a Qdrant path that doesn't exist (Build 32), or would have read an empty local silo ecsmem0d (main 0e7591e).
- Tyrell main 24321a8, a144a1d, fa36509, 9817f7a, 250b8db, 1313eee (Build 33 stamp 1e37a55, tag v0.2.0-build33). Hub tyrelld serves read-only /api/fleet/mem0/{health,v1/telemetry}. Spoke dashboards read it. Spokes run no ecsmem0d, and existing plists are moved aside at onboarding and on every launch. Also fixed two pre-existing test breaks.
- Tyrell Build 33 notarized 82044f4c, stapled, DR unchanged, universal2. Canary rdmbair15m5, then fleet, then hub. Measured app 33 + running daemon-45bc0b545e92 on 5/5 reachable hosts.
- Spoke silos retired 00:21 on rdmpw3265m, rdmpw3275m, rdmbair15m5, jdmbair13m5. launchctl bootout, plist moved to ~/Library/Application Support/Tyrell/retired-launchagents/, data untouched.
- ECSMemory main b067fd9 + bc47aff: exact-text embedding LRU (ECSMEM_EMBED_CACHE_SIZE=512), keep_alive 24h, and cache telemetry. 329/329 tests. Deployed to the hub as pinned releases under ~/Library/Application Support/ECSMemory/releases/b067fd9cbbeb and Tyrell/releases/ecsmem0d-b067fd9cbbeb; ecsmem0d kickstarted.
- Repos: sweep of 206 repos: 162 clean/synced, 5 fast-forward pushes, upstream tracking set on 36 (fleet/overnight/repo-upstream-set-20260926.log), 2 stale 40 h index.lock moved to ~/dev/_quarantine/stale-index-locks-20260926/, and uncommitted docs/todo work snapshotted (~/.agent-coordination/snapshots/rdmsm4x/{docs,todo}/20260926-002006).
- Old Tyrell copies: 2 empty root-owned rollback dirs moved to ~/archive-quarantine-20260926/app-revisions/ (rdmsm4x, rdmbair15m5) with MANIFEST.tsv.
Verification
- Tyrell swift test on rdmsm4x: 1019 Swift Testing / 0 failures (Build 32 gate 984); XCTest 423 (+184, +11), 0 failures. The negative control fails as expected.
- ECSMemory: 212 unit + 117 E2E = 329, 0 failures.
- Relay from rdmpw3265m returns vector_count 27, commit b067fd9, relay tyrelld; /v1/memories → 404 (not relayed).
- Canary 12-min watch: same pid, 30–79 MiB, 24/24 port probes 200.
- Not verified: a visual screenshot of the spoke dashboard.
Undo
- Tyrell: /Applications/.tyrell-rollbacks/<20260926-…>/Tyrell.app (Build 32) per host via install_app_local.zsh, plus the previous daemon release under ~/Library/Application Support/Tyrell/releases/daemon-098e1b457bc4.
- ECSMemory hub: targets in ~/dev/fleet/overnight/ecsmemory-deploy-rollback-20260926.txt (ln -sfn), then kickstart com.eastcoastscience.ecsmem0d.
- Spoke silo: move the plist back from retired-launchagents and
launchctl bootstrap gui/$(id -u)
. Build 33 will retire it again.
Outstanding
- rdmbair13m5: needs someone to check the machine physically. Build 33 install is tracked in TASK-20260926-04.
- Ollama down on both Intel hosts: TASK-20260926-02. GitHub backup gaps (ai/LLM archived, llm-wiki GH007): TASK-20260926-03.
- FEAT-20260924-13 (ecsmem0d over the tailnet) still needs Rich's acceptance of the tailnet exposure. Nothing changed there tonight.