rdmsm4x-changelog-20260830-1230-stability-review-false-failing-fix-and-deploy-directive-triage
Session: claude@rdmsm4x (dev-bf) · 2026-08-30
12:18:05 EDT → 12:30 EDT · host rdmsm4x Task from
Rich: resume all work, coordinate with other agents on this
host, check state and resume accordingly.
Fixed a monitoring defect that was grading the fleet FAILING while every job ran on cadence, and stopped Rich's fleet-wide app deploy directive short of two irreversible mistakes. No app was deployed and nothing was deleted — both by design, on evidence.
1. Coordination (done first, before any write)
Seven live Claude sessions on this host. Scope negotiated with all of them; no collisions.
| Session | Owns | Boundary agreed |
|---|---|---|
dev-76 |
fleet tickets, changelogs | ceded items 1 and 2 to me |
dev-73 |
unscoped at the time | ceded both, will name paths before writing |
dev-d5 |
nothing (fresh) | idle, offered capacity |
dev-e2 |
CPU-storm guard, NEW label
com.eastcoastscience.stormguard |
I must not reap that label; it was blocked on my item 1 |
logtty-d6 |
LogTTY build + sign | I take distribution only, on a hash-pinned artifact |
replicantdb-ca |
replicantDB Sources/ Tests/ dist/ build.sh + git |
I take distribution + /Applications hygiene |
replicantdb-a1 |
replicantDB docs, release-state, fleet verification | supplied the install audit matrix |
Peer claims were treated as unverified until re-measured here. Two turned out to matter and both were confirmed independently before being acted on or escalated.
2.
Defect FIXED — fleet_stability_review.zsh false FAILING
grade (ISSUE-20260830-29, resolved)
The 12:02 review graded rdmsm4x FAILING — agentstatus 0%, agentheal 1%, fleetaudit 2%, fleethealth 5%, devsite 5%, "effectively not running". All five were in fact healthy.
Root cause. Expected runs were computed from host
uptime. A LaunchAgent in the gui/$(id -u) domain does not
exist while there is no Aqua session, so 10 hours with no GUI login were
counted as 10 hours of failure.
Evidence measured here, not relayed:
| Signal | Value |
|---|---|
kern.boottime |
2026-08-30 01:53:16 EDT |
stat -f %m /dev/console |
2026-08-30 12:02:48 EDT |
loginwindow lstart |
2026-08-30 12:01:58 EDT |
| Headless window | 10h 09m with no Aqua session |
Delivery re-measured by run count from session start: agentstatus
(300 s) runs=4 / expected ~4 = 100%; agentheal (600 s)
runs=2 = 100%; fleetaudit (900 s) runs=2 =
100%; fleethealth, devsite, configsync, mcpreconcile,
todoscan, stabilityreview all runs=1, exit 0, state=active. Corroborated
by side effects carrying independent timestamps:
heartbeats/rdmsm4x.json 12:15:30,
config-sync-state 12:10, mcp-sync 12:03,
checkins/todo-scanner 12:09.
Fix — v1.0 → v1.1, file
~/dev/scripts/fleet_stability_review.zsh:
- Window is now
min(uptime, age-since-install, age-since-GUI-session-start). - Session start anchors on
loginwindowlstart, falls back to/dev/consolemtime, clamped to never precede boot nor exceed uptime. - A reboot with no login is reported as explicit context, not a fault.
- New
GUI sessioncolumn; durations humanized so a 21-minute session no longer renders0h.
Validated by re-run over all 6 hosts at 12:23 EDT, not by inspection:
| before (12:02) | after (12:23) | |
|---|---|---|
| Fleet grade | DEGRADED, 68% | MOSTLY STABLE, 91% |
| Degraded hosts | 2 | 0 |
| rdmsm4x | FAILING, 0–5% | stable, 100% on all five jobs |
| rdmbair13m5 | DEGRADED, 70% | stable, 86% (had its own 5 h headless window) |
| rdmbair15m5 | unreachable | unreachable (correctly) |
Why it mattered: the reviewer exists to catch jobs that report success while delivering nothing. Grading them over a window in which they could not run made it commit that same error in reverse — crying wolf on a healthy fleet. Every unattended Mac that reboots overnight would have been graded FAILING, every time.
- Backup:
~/dev/scripts/fleet_stability_review.zsh.bak-20260830-1222 - Undo:
cp ~/dev/scripts/fleet_stability_review.zsh.bak-20260830-1222 ~/dev/scripts/fleet_stability_review.zsh
3. Rich's 11:45 deploy directive — TRIAGED, HELD, nothing shipped
Directive 20260830-114510-F2377EF2 asks to
rebuild/deploy XEntropy, updateRoo, AINetNode; deploy LogTTY,
replicantDB, Tyrell fleet-wide; and remove stale builds. None of
it is ready to ship.
Ship-stopper A — LogTTY bundle identifier (ISSUE-20260830-28,
OPEN, needs Rich). Verified here via PlistBuddy on
Contents/Info.plist, two hosts:
~/dev/apps/LogTTY/dist/logTTY.app -> com.eastcoastscience.logTTY (lowercase l)
rdmsm4x /Applications -> com.eastcoastscience.LogTTY (capital L)
rdmbair13m5 /Applications -> com.eastcoastscience.LogTTY (capital L)
Both universal x86_64+arm64 — architecture is fine, the identifier is
not. To macOS a changed CFBundleIdentifier is a new
application: TCC grants, Keychain items and LaunchServices
lineage do not carry over, and reinstalling the old build does not
restore them. Deploying would have silently deregistered the LogTTY the
fleet has been permissioning for weeks, on five hosts at once. Not
reversible by reinstalling.
It hid because the fleet runs case-insensitive APFS
— /Applications/LogTTY.app and
/Applications/logTTY.app are the same file, so every
filename-based check reports a match. Only Info.plist reveals it.
Corollary adopted for all cleanup: match on bundle identifier,
never on filename.
Ship-stopper B — three apps are arm64-only (ISSUE-20260830-30, OPEN). Measured over ssh with a positive control in the same invocation:
XEntropy / updateRoo / AINetNode -> Mach-O 64-bit executable arm64
rdmpw3265m uname -m -> x86_64
rdmpw3275m uname -m -> x86_64
An arm64-only bundle copies to an Intel host cleanly, passes
existence and checksum checks, and dies at exec with
Bad CPU type in executable. Today nothing is broken — all
three are installed only on the arm64 hosts. The risk is created
by the deploy, not by the status quo. No universal build exists
in the dev tree, so directive item 1 is a build task, not a
deploy task.
replicantDB — no deploy needed. All five reachable
hosts already run 1.10.1 (14) "Quiet", verified by execution (live MCP
handshake returning serverInfo), not by file copy.
replicantdb-ca also found the background daemon has never
run under launchd on any host — the LaunchAgent invokes the app binary
with --daemon, but daemon mode lives in
replicantdb-cli, so it exits 0 and KeepAlive correctly
declines to restart. They are fixing it; I granted the hold rather than
ship 1.10.1 today and 1.10.2 tomorrow to machines that are already
current.
Stale-build cleanup: deferred deliberately. A designated-requirement mismatch between hosts means a future update wipes TCC grants on the odd host out. Nothing is deleted until the install matrix is settled and each removal is matched by bundle identifier.
4. Blocked — needs Rich physically (ISSUE-20260830-25, filed by dev-76)
rdmbair15m5 is up but refuses this fleet's ssh
key. Confirmed here: ICMP replies at 2.9 ms; ssh returns
Permission denied (publickey,password,keyboard-interactive);
control host rdmbair13m5 accepts the same key from the same
session. Fault is on 15m5 — likely clobbered
authorized_keys or permissions too open on ~ /
~/.ssh. Tailscale node offline, WoL did nothing, silent on
the bus since 11:53. Directive item 2 (scan and consolidate 15m5 into
rdmsm4x:~/dev) cannot start until it is
reachable.
5. Unverified-success defect found
agy's 11:53 closeout on the bus claimed it wrote two files.
Neither exists:
~/dev/_handoff/agy-master-closeout-20260830.md and
~/.agent-coordination/handovers/agy-rdmbair15m5-master-closeout-20260830.md.
Also ~/dev/_handoff/staged_LogTTY/ and
staged_replicantDB/ are empty directories from 00:51.
rdmbair15m5 parity must be treated as UNVERIFIED, not
complete.
6. Files touched
| Path | Change |
|---|---|
~/dev/scripts/fleet_stability_review.zsh |
v1.0 → v1.1 (the fix) |
~/dev/scripts/fleet_stability_review.zsh.bak-20260830-1222 |
new backup |
~/dev/fleet/stability/stability-20260830-{1222,1223}.md |
validation runs |
~/dev/issues/ |
ISSUE-20260830-28, -29 (resolved), -30 |
No app binary was deployed. No file was deleted. No launchd job was modified.
7. Outstanding owner actions for Rich
- Decide the LogTTY identifier (ISSUE-20260830-28).
Recommended: keep
com.eastcoastscience.LogTTYcanonical — it is what is deployed and permissioned — and correct the dev tree. Cost: one Info.plist change. The alternative costs re-granting TCC on six machines for no benefit. - Get to rdmbair15m5's keyboard (ISSUE-20260830-25) — ssh auth is broken and it cannot be fixed remotely.
- Assign a universal2 build owner for XEntropy, updateRoo, AINetNode (ISSUE-20260830-30).
No secrets, tokens or credentials appear in this record.
ADDENDUM — 12:29 EDT · two findings that arrived after the report above
A. Git push footgun — do not run a push sweep (ISSUE-20260830-31, dev-76)
Reported by dev-76, verified here
independently by git config, read-only:
apps/replicantDB branch.main.remote = roodb-legacy -> .../rdmsm4x-apps-roodb.git
apps/rooDB branch.main.remote = origin -> .../rdmsm4x-apps-roodb.git
Same URL. A bare git push from
replicantDB writes 125 replicantDB commits into the retired rooDB
repository. Both repos carry a correct backup remote that
neither branch tracks. Same shape on Xentropy, scanRoo, rdLLM,
rdWORKBENCH; logTTY has no push remote at all.
The correction that matters more (dev-76's, and it cuts the
other way): a naive audit reads "277 unpushed commits
across 26 repos" and concludes the fleet is one panic from losing
everything. It isn't. Measured against the canonical backup
remotes with read-only ls-remote, true exposure is
~23 commits, not 277 — replicantDB is 1 behind backup,
not 125. The frightening number is an artifact of the broken upstream,
and "277 unpushed after a kernel panic" is exactly the finding
that triggers an emergency mass-push — the one action that fires the
footgun. Measure against backup/main explicitly,
never @{u}, until 31 is fixed.
B. Correction — rdmbair15m5 is NOT missing from fleet-macs.json
dev-e2 filed the failed Wake-on-LAN as a missing MAC
record. That premise is wrong. Measured:
~/.agent-coordination/fleet-macs.json already holds
rdmbair15m5 en0 fc:b2:14:41:58:bd — the same MAC the wake
was already attempted against. arp -a shows 192.168.1.48
answering as fa:8a:00:50:6d:01, a
locally-administered address (sleep proxy, or
randomized Wi-Fi MAC), so the magic packet went to an address nothing
listens on. No edit to fleet-macs.json is needed. Immaterial to the
outcome either way: 15m5 refuses this fleet's ssh key even when awake,
so it needs Rich at the keyboard regardless.
The pattern across today
Six findings, one defect class — a status that looks like a measurement and isn't:
- Stability reporter grading expected runs against uptime → false FAILING.
- Empty ack asserting acceptance on hard-blocked work, to an unreachable host.
- Filename and checksum checks passing over a changed
CFBundleIdentifier. - Existence and checksum checks passing over an arm64-only binary.
@{u}reporting a 277-commit gap against the wrong repository.- (
dev-d5) A fixture guard asking "did anything change" against a store many agents write concurrently — it caught another agent's ticket instead of its own. On shared state the answerable question is "did anything I would have created appear".
ADDENDUM 2 — 12:33 EDT · replicantDB v1.11.0 is a real deploy candidate
replicantdb-a1 retracted its earlier "no deploy needed"
— correctly, and before I acted on it. It was true at 12:22 and stopped
being true at 12:30.
v1.11.0 (15) "Resident" landed on rdmsm4x at 12:31:05 EDT (commit d80626e, 424 tests, 0 failures). Fleet is on v1.10.1 (14).
Distribution-readiness checks — all pass. Build confirmed settled first (exec mtime and recent-file count stable across two samples 6 s apart), so this is not a mid-write artifact:
| Check | Result |
|---|---|
lipo -archs |
x86_64 arm64 — universal, both Intel hosts safe |
codesign --verify --strict |
exit 0 |
| TeamIdentifier | ZU2882L4HT, DR leaf "Apple Development: Richard Doty (S65Q255HA8)" |
| DR read method | unanchored grep 'designated =>' — a
real DR string, not the empty-output artifact |
| CFBundleIdentifier vs deployed | com.eastcoastscience.replicantDB identical on
all five hosts, all at build 14 |
That last row is the check that caught LogTTY, applied here and passing. No identity change, no TCC or Keychain loss.
Still not shipped. It is mechanically the first genuinely shippable artifact of the day, but it is a behaviour change, not a version bump: the fix makes a daemon that has never run on any host start real work on a ~61k-row live corpus. Recommendation to Rich is rdmsm4x first, observe an actual maintenance pass, then fan out. Go/no-go is his.
The unverified claim, carried verbatim rather than smoothed
away: the verification is against a temp database. It
proves the daemon starts and stays resident — alive past 18 s,
daemon.lock created for the first time, second agent exits
0 leaving the first running. It does not prove the
maintenance pass does useful work on production. "Daemon fixed"
is not "indexing working", and nobody has verified the
latter.
Fleet-wide consequence, previously unknown: the
LaunchAgent invoked the app binary with --daemon, but
daemon mode lives in replicantdb-cli. The app took the
single-instance role, lost it to the menu-bar app, and exited 0 — which
KeepAlive: SuccessfulExit=false correctly declined to
retry. So the 300 s maintenance pass and FSEvents monitoring
have never run under launchd on any host. Every host
has had a registered daemon doing nothing. "Installed and current"
was never the same as "running" — which reframes what the 11:45
directive was actually asking for, and is instance #7 of the day's
defect class.
ADDENDUM 3 — 12:36 EDT · Time Machine protects ZERO fleet hosts (ISSUE-20260830-36)
dev-e2 found rdmsm4x's Time Machine stalled and asked
whether the exposure was fleet-wide. I took the check. It is — and in a
worse form than "they share a broken target".
Measured 12:35–12:36 EDT via tmutil, every remote call
carrying a positive control in the same invocation:
| Host | Destination | Reality |
|---|---|---|
rdmsm4x |
Network, smb://rich@unaspro818a/timemachine |
stalled — 14 MB of 3.9 TB, 0.0136% |
rdmpw3265m |
Local, "Macintosh HD" | no backup at all — mount fails, Code=17 (corrected, see below) |
rdmpw3275m |
Local, "Macintosh HD" | no backup at all — mount fails, Code=17 (corrected, see below) |
rdmbair13m5 |
No destinations configured | BackupNotRunning |
jdmbair13m5 |
No destinations configured | BackupNotRunning |
rdmbair15m5 |
unreachable | ISSUE-20260830-25 |
Two hosts have no Time Machine at all. Two have a destination that is not a backup in any useful sense — a local destination on the same physical disk survives neither disk failure nor loss of the machine, and reports healthy forever while doing so. The one host with a genuine off-machine target is stalled. That is the day's defect class in its purest form: a backup that passes every check by being configured, and protects against nothing.
Scope limit, stated so nobody over-reads it:
~/dev is separately protected by
com.eastcoastscience.devbackup and the per-repo git
backup remotes. This is not "all data is unprotected". It
is "a lost or failed Mac cannot be restored, only
rebuilt." Worth pairing with ISSUE-20260830-31 though — the git
backup remotes are themselves mis-tracked, so both safety nets
are degraded at the same time.
Correction to my own contribution: I attempted a
second independent byte-delta at 20 s and 25 s; my awk
failed to parse the tmutil field both times and returned
empty. The "0 bytes across a full 5-minute pass" figure is
dev-e2's measurement, not a second confirmation. What I
confirmed independently is the stalled state.
This also explains the panic, and my own blind spot.
load1 of 60–100 against 71–97% idle CPU was never CPU
demand — threads blocked on a stuck SMB mount count toward load average
on macOS. A kernel thread wedged behind stalled network I/O is a far
more plausible route to watchdogd missing check-ins for 93
s (panic 01:51:46, reboot 01:53:16) than any userspace CPU consumer.
Consequence for the v1.1 fix I shipped today: a delivery grade
computed from run counts reads GREEN on this host right up until it
panics again, because the jobs genuinely are running. Delivery
and saturation are independent signals and must not be welded into one
report.
Confound flagged to the replicantDB canary
replicantdb-ca has deployed v1.11.0 to rdmsm4x only, as
a deliberate canary, and is sampling an unexplained RSS climb (220 MB →
627 MB in under 3 min, not levelled off) for an hour. rdmsm4x is
a compromised canary right now — wedged mds_stores
and a stalled filesystem-level copy. A daemon doing FSEvents monitoring
and a first index pass on a degraded I/O path may not behave
representatively. An alarming curve should read as "needs a clean
host to confirm", a reassuring one as weak evidence rather than a
clearance. The confound cuts both ways and is now recorded.
ADDENDUM 4 — 12:40 EDT · CORRECTION to Addendum 3, and the panic mechanism
dev-e2 challenged my Time Machine framing. They were
right; I re-verified both points myself before amending
ISSUE-20260830-36.
My framing was wrong in the misleading direction
Measured 12:39 EDT, positive control in each invocation:
rdmpw3265m tmutil latestbackup -> Error Domain=com.apple.backupd.ErrorDomain Code=17
"Failed to mount destination."
rdmpw3275m identical Code=17
Destination registered, ID present, mount fails. So those two hosts do not have "a backup that protects against nothing" — they have no backup at all. My wording was worse than imprecise: it hands the reader a mental model with a restore point in it, inviting the question "is the disk likely to fail?" when the correct question is "there is nothing to restore from." Withdrawn, along with the same-disk claim, which neither of us could confirm and which is moot given Code=17.
Revised count: four of five reachable hosts have ZERO machine-level recovery — rdmpw3265m, rdmpw3275m, rdmbair13m5, jdmbair13m5. rdmsm4x is the only host with any restore point.
This is instance #8, and it was mine. I checked that
a destination was configured and reported as though I had
checked it worked. tmutil destinationinfo is a
presence check; tmutil latestbackup is the delivery check.
I ran the first and reported the second.
Panic mechanism — each link measured, none inferred
Verified here: last successful backup
2026-08-30-012119.backup (01:21:19), newest in
tmutil listbackups; boot 01:53:16; watchdog panic
01:51:46.
01:21:19 backup completes successfully
~01:21+ next backup starts and stalls on the SMB mount
01:51:46 watchdogd misses check-ins for 93 s -> panic, ~30 min into the stall
01:53:16 reboot
Stalled SMB copy →
backupd/diskimagesiod/mds_stores/FPCKService
pile into uninterruptible wait → load1 climbs to 60–100 while CPU is
71–97% idle (blocked threads count toward load on
macOS, so this was never CPU demand) → watchdogd
starves.
Two independent problems sharing a subsystem — do not merge them
dev-e2's point, and the one I'd have got wrong alone:
rdmsm4x is the only host WITH a working recent backup, and it is
the one that panicked. The four hosts with no backup are not
exposed to this failure mode at all. The NAS stall is an
availability problem on one host; the absent backups are a
recoverability problem on four. Fixing the NAS does nothing for
the four; configuring the four does nothing for the panic. Merging them
yields one plausible remediation that half-solves both. They are
separate owner actions.
ADDENDUM 5 — 12:45 EDT · Foundation Models standard verified and corrected (DOC-20260830-01)
replicantdb-ca compiled the app-baseline standard's
[UNVERIFIED] Foundation Models shapes against the real SDK
and routed the dev-wide library work here rather than editing it
from replicantDB's context — correct routing. I
independently recompiled and re-ran their probe before lifting
any marker, because lifting [UNVERIFIED] on a
report rather than on a compile would reproduce
exactly the failure the marker exists to prevent.
Reproduced on rdmsm4x — macOS 27.0 (26A5421a), Xcode-beta
MacOSX27.0.sdk, Swift 6.4. Probe committed at
~/dev/apps/replicantDB/scripts/probe_foundation_models.swift:
clean compile, no diagnostics
on-device SystemLanguageModel: AVAILABLE
PrivateCloudCompute: AVAILABLE
contentTagging use case: AVAILABLE
guided generation OK in 1.05s: kind=invoice confidence=95 (their run: 1.19s)
SystemLanguageModel.tokenCount(sample) = 42
The finding: a disclosure boundary hidden behind a default argument
PrivateCloudComputeLanguageModel is
public, publicly constructible via
init(), conforms to LanguageModel,
and reports AVAILABLE.
LanguageModelSession accepts
model: some LanguageModel. The difference between
on-device inference and sending a user's file contents to Apple's
servers is one argument, at one call site — no entitlement
gate, nothing in the type system, and the safe and unsafe calls
look identical at the point of use and in review.
For replicantDB, DataRoo, ReceiptRoo and bookmarkROO this is a disclosure boundary, not a performance tradeoff. Off-device inference may be right for some feature one day; it is never right by accident, and it is Rich's decision, not a default argument's.
Two API-shape corrections — the standard was wrong; code written against it would not compile
tokenCount(for:)is onSystemLanguageModel(macOS 26.4+), notLanguageModelSession.- No static context-window constant exists. The
standard's guessed number was deleted, not corrected —
measure with
tokenCountand handle the overflow error, which carriescontextSizeandtokenCountso you can report exactly what did not fit.
Changes
to
~/dev/lib/app-baseline/PLATFORM_INTELLIGENCE_AND_PERFORMANCE.md
Backup .bak-20260830-1244. Markers lifted on what was
compiled. New BINDING section "The disclosure
boundary", placed before "Where it genuinely helps",
showing the safe and unsafe call sites side by side. Three binding
rules: bind SystemLanguageModel.default explicitly at every
session construction · ship a test that fails if a PCC model can reach a
session · off-device inference is an explicit, reviewed, user-visible
decision. New section "Capacity" recording 1.05–1.19
s/item ≈ 20.7 hours single-threaded over 62,706 items, so
nobody scopes on-device classification as a corpus-wide sweep.
Deliberately still unverified: "MLX has no ANE target" elsewhere in the standard was not probed and must not be upgraded on the strength of this work. The probing session declined to confirm what it had not measured while everything around it was being confirmed — the most copyable behaviour of the day.
Open, needs an owner: audit every app that
constructs a LanguageModelSession for an explicitly bound
model (none audited yet — the standard now requires it, nothing enforces
it), and write the failing test as a shared fixture in
~/dev/lib so apps inherit rather than reimplement it.
Instance #9, and the subtlest: an API that exists, constructs, and reports AVAILABLE — where availability was never the question anyone should have been asking.
ADDENDUM 6 — 14:09 EDT · rdmbair15m5 returned. Directive item 2 CLOSED, and the backup picture completed.
The host came back — and neither of our diagnoses was right
rdmbair15m5 became reachable at ~14:05 EDT. dev-76 then
pulled the timeline that settles it:
12:01:32 boot
12:08–12:29 SSH DENIED from all four hosts
12:49 richh logs in AT THE CONSOLE
14:05 SSH works — authorized_keys mtime 2026-08-21, UNCHANGED, perms correct, our key present
On a FileVault Mac, sshd cannot read a user's
authorized_keys until that user has logged in at the
console since boot. It still offers publickey and still returns
Permission denied (publickey,password,keyboard-interactive)
— indistinguishable from a bad key. Both of us diagnosed clobbered keys
or bad permissions; all four particulars were false,
and the remediation would have changed nothing while SSH came back
anyway — which would have "confirmed" the wrong cause. FLEET.md
§1 already documented this failure mode and we both reached past it for
a more interesting hypothesis. The lesson is check the
documented failure mode before diagnosing a novel one, not "we
misread the error". No action on key material, ever, for this;
sudo fdesetup authrestart is the planned-reboot answer.
Directive item 2 — satisfied by measurement, not by consolidating (TASK-20260830-01, resolved)
Read-only assessment, 14:06–14:09. Nothing pulled, moved,
promoted or deleted. rdmbair15m5:~/dev = 18 GB,
191,759 files. It holds no unique work.
Exactly two repos carried anything not on a remote, and both counts were misleading — resolved by direct SHA comparison rather than trusted:
| Repo | Raw count said | Truth |
|---|---|---|
net/cloudflare-access-passkeys |
6 unpushed | All 6 SHAs present on rdmsm4x — identical history |
apps/LogTTY |
1 unpushed | Stale build-18 snapshot vs hub's 111 commits at build 21 — do not merge |
~/dev/rdmbair15m5 — the directive's second path — is an
empty directory. Half the target did not exist.
Why the counts weren't trusted:
git log --branches --not --remotes compares against
whatever remotes are configured, and ISSUE-20260830-31
established those are mis-tracked. A count from it can be right about
the number and wrong about the meaning. Both were.
This supersedes agy's 11:53 closeout, which asserted parity complete and cited two files that do not exist. Its conclusion was right; its evidence was not.
The backup picture, completed — and a new single point of loss
rdmbair15m5 tmutil latestbackup -> "Failed to find any backups found for current machine"
Configured against the correct NAS, never completed a backup. Final count, all six measured: five of six hosts have ZERO machine-level recovery. rdmsm4x is the only host with any restore point, and its current backup is the stalled one that preceded the panic.
Compounding:
net/cloudflare-access-passkeys — 6 commits, identical on
rdmsm4x and rdmbair15m5, no git remote on either host,
and no working Time Machine on either. That project exists on
two local disks and nowhere else. It is not covered by
ISSUE-20260830-31, which concerns repos with mis-tracked
remotes; this one has none at all.
FLEET.md §1 — measured hardware inventory added
Backup .bak-20260830-1406. RAM, disk and macOS build for
all six hosts, measured. Surfaced two things:
jdmbair13m5 has 135 Gi free of 926 Gi — an
order of magnitude less headroom than any other host, so nothing should
be promoted there — and the table's macOS column for it was
stale (said 26.6.2, actually 27.0). The two Intel hosts
are on 26.7 while all four Apple Silicon hosts are on 27.0.
ADDENDUM 7 — 20:44 EDT · Rich transferred LogTTY ownership. A shipping trap, and a rule conflict.
Rich, 20:38 EDT: "resume all - make claude@rdmsm4x take ownership of LogTTY if that is the question - approved for all". Taken.
🛑
dist/logTTY.app is a trap — broadcast as do-not-ship (bus
20260830-204140-4F31D7BF)
It was rebuilt at 14:51 today, after this morning's measurement, and is a regression in every dimension at once:
| This morning (12:25) | Now (14:51 rebuild) | |
|---|---|---|
| Build | 20 | 18 — fleet runs 19, HEAD intends 21 → downgrade |
| Arch | universal x86_64 arm64 |
arm64 only → dies on both Intel hosts |
codesign --verify --strict |
clean | "a sealed resource is missing or invalid" |
| Signature | Team ZU2882L4HT |
ad-hoc, no team → TCC re-prompts every launch |
Root cause: BUILD_NUMBER is reverted
21 → 18 as an uncommitted working-tree change (HEAD
23f9a89 is literally "bump BUILD_NUMBER to 21"),
and the build was packaged from that dirty tree. Same class as the 08-29
23:09 sweep that left replicantDB unsigned.
Checked before alarming anyone: the repo is INTACT —
111 commits, HEAD 23f9a89, no reset in the reflog, and
3d3c908 is a legitimate ancestor rather than spoke content
pulled backwards. Known-good reference is
/Applications/LogTTY.app on rdmsm4x: build 19,
universal, strict-clean, ZU2882L4HT. Not
dist/.
Identifier decision — settled by the house rule, not by my judgement
~/dev/apps/CLAUDE.md §5: "Bundle IDs:
com.eastcoastscience.<Product>." Surveying
where the identifier is set rather than mentioned, the repo is
already overwhelmingly capital-L — build_and_run.sh (which
packages the app), verify_product_identity.sh,
deploy_release_remote.sh, the packaging test, the CloudKit
string, the mobile target. Lowercase survives in one
subsystem (the privileged helper).
Decision: com.eastcoastscience.LogTTY.
The product name stays "logTTY" — the rename was a
product rename and identifiers are opaque plumbing users never
see, so keeping capital-L costs it nothing. No host loses TCC or
Keychain. The helper subsystem reconciles to it.
My own correction, twice over: I first recommended
capital-L, then withdrew it on finding the rename was deliberate, then
restored it on finding where the identifier is actually set. And I told
a peer I could not find the CloudKit rule in apps/CLAUDE.md
— it is there, in §5; I searched §3 and concluded from its
absence. That is the day's own defect class, committed by me,
for the second time.
PLATFORM_TARGETS.md — a real conflict, reconciled rather than overwritten
Rich set a new standing rule (via replicantdb-a1):
universal2 is a build gate, lipo -archs
fails the build. But PLATFORM_TARGETS.md already carried
Rich's 2026-08-22 decision that Intel ships as "a
separate scheme with its own deployment target, not a
universal binary". A blanket universal mandate reads as overturning
that, and silently overwriting one recorded decision of Rich's with a
later relayed one was not something to do quietly.
The two rules answer different questions, which is what dissolves it:
| Question | Answered by |
|---|---|
| Does the binary contain code this CPU can execute? | the lipo gate |
| Will this OS version load it at all? | the 26.7 scheme |
Intel support requires BOTH — an app with a 27.0
deployment target cannot run on Intel even with an x86_64
slice, because no Intel Mac will ever run 27.x. So
lipo showing two slices is not evidence of
Intel support. Both rules stand; neither is overturned. Written into the
file as an explicit "read both" table. Backup
.bak-20260830-2043. Host table also refreshed — builds had
moved and jdmbair13m5 was missing entirely.
ADDENDUM 8 — 20:52 EDT · CloudKit: the real blocker, named (SEC-20260830-04)
Rich, 20:47: "I've been asking for cloudkit for weeks." Root
cause found. It is not what apps/CLAUDE.md §3 has
been telling every session for weeks.
Measured
— every provisioning profile on rdmsm4x, decoded with
security cms -D
| Profile | App ID | Container | Platform |
|---|---|---|---|
net.dataroo.RTTy.mobile |
ZU2882L4HT.net.dataroo.RTTy.mobile |
iCloud.net.dataroo.RTTy |
iOS |
com.eastcoastscience.Tyrell |
ZU2882L4HT.com.eastcoastscience.Tyrell |
iCloud.com.eastcoastscience.Tyrell |
iOS |
| team wildcard | ZU2882L4HT.* |
none | iOS |
Three profiles, valid into 2027. All three iOS. There is not one macOS provisioning profile on this machine, and no CloudKit container exists for any macOS app. Every app on this fleet is a Mac app. The transports have been ready for weeks — 440 + 73 + 218 lines — pointing at containers that were never created.
The trap: the team wildcard
ZU2882L4HT.* carries no containers and structurally
cannot — Apple forbids iCloud containers on a wildcard App ID. So
signing succeeds against it, builds pass, and CloudKit
silently does nothing.
Why every attempt blamed something different: the
layers fail in sequence — no container → no capability on the App ID →
no macOS profile → no entitlement in the app — and each attempt hit
whichever layer it reached first and reported that. §3's
"the blocker has always been schema promotion and signing" is
half right: it is signing, specifically macOS App IDs and
profiles that were never created. Nothing can be
promoted to a schema in a container that does not exist. That
sentence has been routing sessions to the wrong layer; corrected
in apps/CLAUDE.md §3 (backup
.bak-20260830-2051).
Done tonight — everything not needing Rich's Apple ID
- Canonical naming settled:
iCloud.com.eastcoastscience.<Product>, per-app — matching the one working precedent (Tyrell) andreplicantDB/docs/CLOUDKIT-PLAN.md. ~/dev/apps/LogTTY/entitlements/LogTTY.entitlementswritten,plutil -lintclean, capital-L per Rich's instruction. Deliberately not wired intobuild_and_run.shyet: a restricted entitlement with no matching profile yields an app that will not launch. Added as a new untracked file — it does not touch the 45 dirty files.~/dev/fleet/CLOUDKIT-UNBLOCK.md— the ordered portal checklist.
Needs Rich — ~20 minutes at developer.apple.com, LogTTY first to prove the flow
App ID (macOS, explicit, iCloud capability) → create container
iCloud.com.eastcoastscience.LogTTY typed once,
exactly, permanent → assign → macOS App Development profile →
install. Or let Xcode's Automatically manage signing do it once
the container exists. Then replicantDB (its entitlements are already
correct and waiting) and RTTy (adopt a new container; the dead
iCloud.net.dataroo.RTTy cannot be removed).
Verification rule recorded: never report CloudKit working until a record written on one Mac is read back on another. Every layer passes its own check while the layer below is missing — instance #10 of the day's defect class, and the one that cost the most calendar time.
Also this pass
~/.claude/CLAUDE.md — platform targets added to the
global macOS rules at Rich's request: deployment targets, the two Intel
hosts confirmed by Rich on 26.7 (25G224), the universal2
build gate, the necessary-but-not-sufficient trap, and the
bundle-ID/CloudKit/signing rules. Backup
.bak-20260830-2048.
ADDENDUM 9 — 22:04 EDT · I got one wrong. rdmbair15m5 consolidation REOPENED (TASK-20260830-06)
replicantdb-a1 routed an intake item that contradicted a
conclusion I had published. It was right and I was wrong. Corrected in
the ticket rather than quietly.
What I wrote (TASK-20260830-01): "rdmbair15m5
holds NO unique work. Parity is now VERIFIED by measurement."
What I actually measured: no git repo
under rdmbair15m5:~/dev carries commits absent from the
hub. That narrower claim still stands. The gap: I
generalised a repo-scoped instrument into a claim about all
work. My scan walked
find ~/dev -name .git -maxdepth 4, so a non-git
directory was invisible to it by construction, and
~/Archived-migrated-2026-08-13 was never in scope
at all — it is not under ~/dev. Both exclusions
contained real content.
Measured 22:02–22:03, positive control per call
| Path | State |
|---|---|
~/Archived-migrated-2026-08-13/ |
89 GB, 605,011 files — never triaged by any pass |
~/dev/obsidian_fleet_sync |
not a git repo — 234 files; 190 tests; its own doc
names a canonical home ~/dev/agy/obsidian_fleet_sync/
that does not exist |
~/dev/fleet-agent-registry |
not a git repo — 15 files, 92 MB |
| ticket store | 378 tickets vs this host's ~113 — needs the
ticket CLI, not a file merge |
a1 reported the archive as 48 GB; measured 89
GB. Correcting upward — the finding is theirs.
Relay manifest verified independently:
HANDOFF.md, 7,707 bytes, shasum -a 256 =
b7bce9044fd9f0921423f4000808969089d8b2d54ff35665a5572cbd60994f1d
— exact match to what codex@rdmbair15m5 published on both
ends. Recorded RECEIVED, not accepted: checksum parity
proves the file crossed intact and nothing about
whether its content is correct or complete.
Sequence — deliberately not starting with the 89 GB
obsidian_fleet_sync first (smallest, highest confidence,
closes a doc that already claims a home that does not exist) →
fleet-agent-registry but identify the binaries
first: 92 MB in 15 files usually means a database or cache,
which should not become canonical → Desktop HTMLs, almost certainly
regenerable exports → ticket store, CLI only → the archive LAST
and by sampling. Establish what fraction is unique
before any transfer; it predates the canonical-root discipline
and is likely mostly duplicates. And nothing from it should land
anywhere reaching jdmbair13m5, which has 135 Gi
free of 926 Gi.
Instance #11, and this one is mine
Same shape as the error I flagged in someone else's Time Machine
premise four hours earlier, and as my own
tmutil destinationinfo report: I confirmed the
narrow technical fact and reported the broad operational
conclusion. The rule now recorded fleet-wide — "a search
that matches nothing is not evidence of absence; confirm the search
could have found it" — applies equally to a search that matches
everything it can see. A repo scan cannot see a
non-repo. State the scope of the instrument in the conclusion, every
time.
ADDENDUM 10 — 22:09 EDT · Item (a) closed by verification, not by promotion
I was one command from creating a third copy of a project that was already fully promoted.
Measured: ~/dev/apps/ObsidianFleetSync
is a git repo, 2 commits, 241 files — and the file-set difference
against rdmbair15m5:~/dev/obsidian_fleet_sync is
zero. Not zero-excluding-scratch; zero, including every
.agents/ working file. The hub copy is a strict
superset of the spoke copy. Promotion was completed 2026-08-29
under ISSUE-20260829-04, and that ticket was accurate.
Why it convincingly looked unpromoted — four stale artifacts all pointing the same wrong way
- The project's own docs declared "Canonical
home:
rdmsm4x:~/dev/agy/obsidian_fleet_sync/" — a path that does not exist. It went toapps/, notagy/. ~/dev/todo/items/20260826-obsidianfleetsync-never-promoted-no-git.mdstill said never promoted, no git — superseded three days later and never closed.- A partial duplicate at
~/dev/apps/obsidian_fleet_sync(lowercase — a genuinely different inode,342109867vs340193522, not a case-insensitivity artifact): 72 files, not a git repo, and a strict subset. Anyone checking that path finds a non-repo and concludes "not promoted". That is almost certainly what happened. - A fourth copy staged at
~/dev/_inbox/rdmbair15m5/obsidian_fleet_sync.
Every link in that chain is individually true and the conclusion is false. The check that broke it was comparing file sets between hub and spoke rather than checking whether a named path existed. Existence of a path is a presence check; set difference is a delivery check. Instance #12 — and the first one caught before acting rather than after.
Fixed at the source (commit
920dc4e)
The stale canonical-home claim is corrected across
SESSION-STATE.md, ISSUES.md,
RECONCILIATION-20260823.md and
OBSIDIAN-CONSOLIDATION-AND-ROODB-PLAN-20260823.md — because
it was actively harmful, not untidy: it caused one
false "unpromoted work" finding tonight and would have caused more. The
superseded todo item is closed with the evidence.
(My first substitution ate a path separator and produced
~/dev/apps/ObsidianFleetSyncRECONCILIATION-20260823.md.
Caught on the verify pass, restored from backup, redone correctly.
Noting it because the verify pass is the only reason it did not
ship.)
Nothing deleted. The two redundant copies are
provably lossless to remove and are recorded for removal — but a
directory deletion at 22:09 unattended buys nothing when the finding is
already recorded. And explicitly: do NOT create
~/dev/agy/obsidian_fleet_sync/. That is the
trap.
Credit where it is due
replicantdb-a1 raised this in good faith from a
read-only scan and the signal genuinely looked like unpromoted work. The
finding was wrong; raising it was right — it surfaced
four stale artifacts that were actively misleading every agent that
looked, and those are now fixed.
ADDENDUM 11 — 22:12 EDT · Withdrawing my own jdmbair13m5 constraint (Rich reviewed it)
dev-76 walked back an operational rule they had relayed
to Rich. The inference underneath it was mine, and it
is now withdrawn in both places I put it.
What I wrote in FLEET.md §1: jdmbair13m5's 135 Gi free is "the only one where a large sync or index could plausibly run it out." And in TASK-20260830-06 step (e): "nothing from this archive should be promoted anywhere that lands on that host."
Both withdrawn. Rich reviewed it and 134 Gi is fine for that host's role. Two reasons, both sound:
- jdmbair13m5 is a spoke/scratch node. Canonical is
rdmsm4xby definition — so "don't promote canonical there" warned against something the architecture already excludes. - 86% used on APFS is not the 86% the HFS+-era instinct reacts to. APFS does not fragment-degrade near capacity, and local snapshots purge under pressure.
What survives, and is worth keeping: jdmbair13m5 is
the fleet's only 1 TB host (the others are 4 TB or 8
TB) and the only one above 50% used — 1.0 TB
APPLE SSD AP1024Z, 767 Gi used / 134 Gi available, no
external disks. FLEET.md now records that as "inventory, not a
constraint" and marks the prohibition withdrawn in the
file rather than deleting the sentence, so a later reader
cannot re-derive it from a bare number.
A genuinely different route to the same failure
Every other instance today was one agent over-reading its own
instrument — my repo scan, my
tmutil destinationinfo, the anchored designated-requirement
grep. This one only exists because the claim crossed between
agents.
I attached a hedged inference to a measured number, in a file five
other agents read as ground truth. dev-76 relayed it to
Rich as a firm rule without measuring it. Rich challenged it — "4TB
drive - you sure it's not your script?" — and only then did anyone
verify. The number held; the conclusion never had
support.
A verified number does not confer verification on the inference travelling beside it. It arrives as context attached to something that was checked, so it never presents as a claim needing checking. That is why neither of us checked it.
One concrete fix on my side: the 4 TB Rich was thinking of is
rdmbair15m5, which sat two lines above
jdmbair13m5 in my own report. Adjacent rows with similar names and very
different numbers invite misreads — worth guarding against in how
per-host tables are laid out, not just in what they say.