Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260831-1236-replicantdb-daemon-progress-ocr-bounds-and-shared-cloudkit

replicantDB — daemon made to run, indexing bounded, progress surfaced, CloudKit shared

Host: rdmsm4x · Session: claude@rdmsm4x · Window: 2026-08-30 12:20 EDT → 2026-08-31 12:36 EDT (times read from date, not estimated)

One line: the background daemon had never run on any host in the project's history; it now runs, survives files that used to freeze it permanently, bounds every synchronous framework call on the indexing path, reports its own progress live, and indexes a 2.08M-file corpus from a clean database.

Versions shipped

1.10.1 → 1.17.0 (32) "Shared", 27 commits. Fleet went 1.10.1 → 1.16.1 on all six hosts.

The defects that mattered

The daemon never ran. The LaunchAgent invoked the app binary with --daemon; the app binary had no daemon mode (it lived in replicantdb-cli). It exited 0, KeepAlive SuccessfulExit=false correctly declined to restart it, and both log files stayed empty. "Daemon installed" had never once meant "daemon working". Fixed v1.11.0; proven by daemon.lock existing for the first time.

Making it run exposed an autorelease leak. RSS climbed 527→1392 MB in four minutes, 1.1 GB of it in 555 dirty IOSurface regions. No autoreleasepool anywhere in Sources/. Correct under the app's run loop, wrong under dispatchMain(). Fixed v1.11.1–1.11.3; growth fell 81× and memory began being returned, which a leak never does.

One image could wedge the daemon forever. VNImageRequestHandler.perform is synchronous and unbounded. sample caught it: identical stack in all 7708 samples. The walk never returned, so PassGate.leave() never ran, so every subsequent trigger was dropped for the process's life — presenting as a healthy idle daemon. Bounded v1.14.0, generalised to all per-file extraction v1.14.1.

Vision hangs are common, not rare. Twelve hangs in twenty-four minutes of one home directory. The bound turned a wedge into a crawl. Circuit breaker added v1.16.0 — and its first design was wrong: it counted consecutive hangs, which do not occur, so it never fired (16 leaked threads in 4 minutes). v1.16.1 made it a total budget a success does not restore, because the resource is leaked threads and a success does not return one.

replicantdb-cli --daemon crashed nondeterministically. dispatchMain() traps off the main thread; the CLI called it from an async context. Ran correctly on rdmsm4x every time, crashed 3/3 on rdmbair13m5. Fixed v1.13.1.

Delivered on request

Live indexing progress with per-stage bars and auto-refresh (no button); fresh database (old 307 MB / 84,278-row store archived to ~/Library/Application Support/replicantDB/archive-20260830T205905, integrity verified); Location and Local Network permission checks (Location was declared in Info.plist and never probed); System/Light/Dark switcher; macOS 27 storage-API research (docs/research/MACOS-27-STORAGE-ACCESS.md — nothing useful there, recorded so it is not redone); ECSCloudKit adopted and pushed to github.com/richhdoty/ECSCloudKit private.

Verification

461 tests here + 22 in the shared package. Every fix mutation-tested. Deploys now read from dist-snapshots/<version>-build<N>/, never live dist/.

NOT proven

Undo

git revert any commit; nothing pushed from this repo. Old database is archived, not deleted. Previous builds are in dist-snapshots/.


Addendum — 2026-08-31 14:36 EDT: credential containment (ISSUE-20260831-05)

Six repositories were found to hold live Cloudflare, GoDaddy and TLS private key material. This session was asked whether replicantDB's corpus was a separate exposure. It was.

preview_snippet stores up to 2,000 characters of each file's CONTENT, queryable from the CLI, the MCP server and search. Indexing those directories copied their contents into a second store with different readers and a different blast radius than git.

Store Before After
Live corpus rows from the six paths 57,090 0
…carrying content snippets 41,698 0
FTS search-index references — 0
Eight on-disk DB copies (7 backups + 1 archive) 634 each = 5,072 0
Live DB free pages holding deleted bytes 34,588 (135 MB) 0

v1.19.0 "Contained" enforces a denylist at RooDatabase.upsertFile / upsertFilesBatch — not in the indexer — because every row enters through those two functions. purgeDeniedPaths() runs on every daemon start, since a denylist added afterwards does not un-index what is already stored. Default list extends to ~/.secrets, ~/.ssh, ~/.aws, ~/.config/gh.

Verified in production, not only in tests: 146,214 files remain on disk under two denied paths while the daemon actively indexes, and denied-path rows held at 0 across four minutes.

Two errors of mine, recorded because they transfer

NOT remediated — Rich's, and rotation is the real fix

Time Machine is enabled to smb://rich@unaspro818a/timemachine, with nine local APFS snapshots spanning 2026-08-30 15:33 to 2026-08-31 12:42 — the last roughly 100 minutes before the purge. Those snapshots and the NAS hold the pre-remediation state, of the corpus AND of the affected repositories. "The credentials are no longer on disk anywhere" is not a claim anyone can currently make, which is why rotation rather than deletion is the remediation. No snapshot was touched; that is destructive and Rich's call.