rdmsm4x — replicantDB on all six hosts; fleet MAC inventory completed; a third Tyrell host repaired
2026-08-28 15:51:23 EDT · rdmsm4x Fourth entry in
this thread. Supersedes the "5 of 6" state in
…20260828-1545-replicantdb-fleet-whole-and-runbook.
The fleet is complete for the first time
rdmpw3265m returned after roughly two days down and was
caught within two minutes of booting. Three things were done in that
window, in deliberate priority order:
1. Its MAC addresses were captured FIRST. That is
the single action that requires the host to be up, and it had been
blocking the whole fleet's ability to wake a sleeping machine. A
Wake-on-LAN magic packet needs the target's MAC and has no
broadcast form, so until now a sleeping host could only be
recovered by walking to it. All six hosts are now
recorded in ~/.agent-coordination/fleet-macs.json
(mode 600) — deliberately outside ~/dev/data/, which feeds
fleet_status.zsh and thence the dashboard published to
dev.dataroo.net.
2. v1.4.0 deployed and verified by execution, not by
copy: universal2 binary on an x86_64 host,
codesign --verify --strict clean, CLI self-reporting
1.4.0 (5) "Reference", and a pipelined MCP
initialize returning
serverInfo {name: replicantdb, version: 1.4.0}.
All six hosts are now on v1.4.0, both Intel machines included.
3. The cross-check flagged while the host was unreachable was
correct. It carried an arm64-only
tyrell-mcp on an x86_64 machine, and it was
registered in ~/.claude.json — so a configuration check
reported it present and working while it could never execute a single
byte (Bad CPU type in executable). Three of
three affected hosts now confirmed. Repaired with rdmsm4x's
universal binary and handshake-verified
(serverInfo {name: tyrell, version: 0.1.0}).
The underlying defect is still open in
apps/Tyrelland should stay open. Hand-repairing three hosts does not fix a distribution path that ships whatever single-arch artifact the building host happened to produce. The next host added will have the same problem.
Also fixed today: a version that shipped wrong through an entire release
A pipelined MCP handshake against a freshly updated host returned
serverInfo.version = "1.3.0" while the CLI
reported 1.4.0 (5). ReplicantDBMCP/main.swift
hardcoded the string, so every MCP client on five hosts was told the
wrong version for the whole of v1.4.0. Nothing failed and nothing warned
— a second copy of a version does not announce itself, it just
disagrees.
VersionSingleSourceTests now walks Sources/
and fails, naming the offending file, if the literal reappears.
I verified the test genuinely catches it by
reintroducing the hardcode and watching it fail; a guard test never seen
to fail is not evidence of anything.
A second lesson from the same hour: the first fleet
verification pass used sed to pull the version out of the
handshake JSON and silently confirmed only one host of
five, because serverInfo key order is not stable.
It looked like a clean pass. Parse JSON; do not pattern-match it.
Documentation written
UPDATING.md— release and fleet-update runbook. Each step names the defect that justifies it rather than asserting ceremony, and it carries a diagnosis table that distinguishes a powered-off host (no ARP from several peers) from a FileVault-locked one (ping and sshd fine, host key matches, key rejected) from a reimaged one (host key differs).RESUME.md— rewritten as a cold-start entry point around the most useful fact about this codebase: every defect that has mattered was found by running the product, and several were invisible to a green 300+ suite. Listed beside what the suite said at the time.HANDOFFS.md,SESSION-STATE.md,ISSUES.md,CHANGELOG.md,~/dev/PROJECTS.mdand the dashboard fact source all current.
State
v1.4.0 (5) "Reference" · 340 tests, 0 failures, 0 warnings · tree clean · 6 of 6 hosts. Tags v1.1.0 · v1.2.0 · v1.3.0 · v1.4.0. Nothing pushed to any remote — D-26 remains Rich's call; the git history still carries employer client names, so a clean working tree is not a publishable repo.
Next — a decision before code
The classifier residual is images classified from OCR
prose: 19 png receipts, 17 png
contracts, and a photograph filed as a credential. Whether a
photo of a receipt IS a receipt is a design question, not a
defect — both answers are defensible and the wrong one is
expensive to reverse across 61k rows. It belongs in
DECISIONS.md first.
Waiting on Rich: Full Disk Access click · CloudKit container name (team-permanent) · D-26 before anything is pushed · daemon registration (D-4) · and one small new one: confirm Energy → "Wake for network access" on both Intel desktops, since a recorded MAC is useless if the NIC is not listening while asleep.