Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260827-1457-replicantdb-v1.4.0-reference

rdmsm4x — replicantDB v1.4.0 "Reference" shipped and applied

2026-08-27 14:57:13 EDT · rdmsm4x · ~/dev/apps/replicantDB Continues the session recorded in rdmsm4x-changelog-20260826-1910-replicantdb-releases-and-handoffs.md.

Shipped

v1.4.0 "Reference" — tag v1.4.0, commit 4037586. 338 tests, 0 failures, 0 warnings. Universal2 (x86_64 arm64), signed TeamIdentifier=ZU2882L4HT, Metadata.appintents with 29 intent references. Nothing pushed to any remote — D-26 is Rich's call alone.

FilePurpose.reference decides documentation structurally — documentation directory component or documentation filename stem — before any prose is consulted (D-30).

Applied to the live corpus

purpose before after
credential 172 57
report 209 146
correspondence 66 38
receipt 56 42
contract 38 35
other 25,899 24,465
reference 0 1,657

1,657 rows changed. integrity_check = ok after; the single USER_OVERRIDE untouched; an APFS-clone backup taken and integrity-verified before the write.

The correspondence and report losses were sampled by hand before applying, because those are exactly where a wrong rule would do damage — every one is a SKILL.md, a references/*/gotchas.md, a PRD or a master spec. 38 correspondence and 146 report rows survive, so neither type was emptied. Real financial documents intact: paypro_paypal_129.85_070226.png → invoice, tumi_ro-20260618534718.pdf → receipt.

The design constraint that made this a new type, not a stricter tier

A plain .md carrying Subject:/From: headers is still correspondence. That negative regression test is the point of the release: without it, adding a type and tightening the archetype tier are indistinguishable from the outside, and only one of them is correct.

A crash found by running it, not by testing it

The new code built the filename stem above the bookmark early-return, so it had to tolerate every row in the table — and 1,260 live rows are browser bookmarks whose filename IS an http(s) URL. On this Foundation, URL(fileURLWithPath:) traps with API MISUSE: URL(filePath:) called with an HTTP URL string rather than returning nil. The whole classify pass died on the first such row. Now pure string handling, with two regression tests. No fixture among 336 tests was shaped like a bookmark URL.

Delegation, and why the review mattered

One codex exec --model gpt-5.6-sol job (OpenAI at 16%, proceed; Anthropic weekly at 91%, hold). It reported 23 test failures it believed environmental — and it was right: the same suite ran 0 failures here, the sixth confirmation of that sandbox pattern. It also correctly refused to edit those tests to pass.

Its five legacy-test edits were correct on review: every not: assertion untouched, only the is: expectation moved, and each test renamed so its name stayed truthful. But the crash above was in its code and its 336 tests did not catch it — the diff, not the summary, is what review means.

Deployed 4 of 6 — and the two absences are different failures

rdmsm4x, rdmbair13m5, rdmpw3275m (Intel — CLI executed there), jdmbair13m5.

The generalisable fix, filed to the fleet: a scripted reboot of a FileVault Mac must use sudo fdesetup authrestart, never plain reboot — otherwise a remote maintenance run locks itself out of the machine it was maintaining. Worth auditing update_mac.zsh.

NOT proven — the residual changed shape rather than closing

Remaining false positives are images classified from OCR prose: 19 png receipts, 17 png contracts, 17 png + 10 jpeg credentials — one is oops_photos/IMG_3805.jpeg, a photograph. Whether a photo of a receipt IS a receipt is a genuine design question, not a defect; it goes to DECISIONS.md before any code, because both answers are defensible and the wrong one is expensive to reverse across 61k rows.

Two cheap non-design follow-ups: reports is missing from documentationDirectoryComponents (while references is present), and 15 .md stems are outside the vocabulary.

Still unproven from before: Siri has never been asked to invoke an intent; the ≤10-of-64 perceptual threshold is untuned against ground truth; --apply dedup has never run.