Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260830-1230-stability-review-false-failing-fix-and-deploy-directive-triage

rdmsm4x-changelog-20260830-1230-stability-review-false-failing-fix-and-deploy-directive-triage

Session: claude@rdmsm4x (dev-bf) · 2026-08-30 12:18:05 EDT → 12:30 EDT · host rdmsm4x Task from Rich: resume all work, coordinate with other agents on this host, check state and resume accordingly.

Fixed a monitoring defect that was grading the fleet FAILING while every job ran on cadence, and stopped Rich's fleet-wide app deploy directive short of two irreversible mistakes. No app was deployed and nothing was deleted — both by design, on evidence.


1. Coordination (done first, before any write)

Seven live Claude sessions on this host. Scope negotiated with all of them; no collisions.

Session Owns Boundary agreed
dev-76 fleet tickets, changelogs ceded items 1 and 2 to me
dev-73 unscoped at the time ceded both, will name paths before writing
dev-d5 nothing (fresh) idle, offered capacity
dev-e2 CPU-storm guard, NEW label com.eastcoastscience.stormguard I must not reap that label; it was blocked on my item 1
logtty-d6 LogTTY build + sign I take distribution only, on a hash-pinned artifact
replicantdb-ca replicantDB Sources/ Tests/ dist/ build.sh + git I take distribution + /Applications hygiene
replicantdb-a1 replicantDB docs, release-state, fleet verification supplied the install audit matrix

Peer claims were treated as unverified until re-measured here. Two turned out to matter and both were confirmed independently before being acted on or escalated.

2. Defect FIXED — fleet_stability_review.zsh false FAILING grade (ISSUE-20260830-29, resolved)

The 12:02 review graded rdmsm4x FAILING — agentstatus 0%, agentheal 1%, fleetaudit 2%, fleethealth 5%, devsite 5%, "effectively not running". All five were in fact healthy.

Root cause. Expected runs were computed from host uptime. A LaunchAgent in the gui/$(id -u) domain does not exist while there is no Aqua session, so 10 hours with no GUI login were counted as 10 hours of failure.

Evidence measured here, not relayed:

Signal Value
kern.boottime 2026-08-30 01:53:16 EDT
stat -f %m /dev/console 2026-08-30 12:02:48 EDT
loginwindow lstart 2026-08-30 12:01:58 EDT
Headless window 10h 09m with no Aqua session

Delivery re-measured by run count from session start: agentstatus (300 s) runs=4 / expected ~4 = 100%; agentheal (600 s) runs=2 = 100%; fleetaudit (900 s) runs=2 = 100%; fleethealth, devsite, configsync, mcpreconcile, todoscan, stabilityreview all runs=1, exit 0, state=active. Corroborated by side effects carrying independent timestamps: heartbeats/rdmsm4x.json 12:15:30, config-sync-state 12:10, mcp-sync 12:03, checkins/todo-scanner 12:09.

Fix — v1.0 → v1.1, file ~/dev/scripts/fleet_stability_review.zsh:

Validated by re-run over all 6 hosts at 12:23 EDT, not by inspection:

before (12:02) after (12:23)
Fleet grade DEGRADED, 68% MOSTLY STABLE, 91%
Degraded hosts 2 0
rdmsm4x FAILING, 0–5% stable, 100% on all five jobs
rdmbair13m5 DEGRADED, 70% stable, 86% (had its own 5 h headless window)
rdmbair15m5 unreachable unreachable (correctly)

Why it mattered: the reviewer exists to catch jobs that report success while delivering nothing. Grading them over a window in which they could not run made it commit that same error in reverse — crying wolf on a healthy fleet. Every unattended Mac that reboots overnight would have been graded FAILING, every time.

3. Rich's 11:45 deploy directive — TRIAGED, HELD, nothing shipped

Directive 20260830-114510-F2377EF2 asks to rebuild/deploy XEntropy, updateRoo, AINetNode; deploy LogTTY, replicantDB, Tyrell fleet-wide; and remove stale builds. None of it is ready to ship.

Ship-stopper A — LogTTY bundle identifier (ISSUE-20260830-28, OPEN, needs Rich). Verified here via PlistBuddy on Contents/Info.plist, two hosts:

~/dev/apps/LogTTY/dist/logTTY.app  -> com.eastcoastscience.logTTY   (lowercase l)
rdmsm4x     /Applications          -> com.eastcoastscience.LogTTY   (capital L)
rdmbair13m5 /Applications          -> com.eastcoastscience.LogTTY   (capital L)

Both universal x86_64+arm64 — architecture is fine, the identifier is not. To macOS a changed CFBundleIdentifier is a new application: TCC grants, Keychain items and LaunchServices lineage do not carry over, and reinstalling the old build does not restore them. Deploying would have silently deregistered the LogTTY the fleet has been permissioning for weeks, on five hosts at once. Not reversible by reinstalling.

It hid because the fleet runs case-insensitive APFS — /Applications/LogTTY.app and /Applications/logTTY.app are the same file, so every filename-based check reports a match. Only Info.plist reveals it. Corollary adopted for all cleanup: match on bundle identifier, never on filename.

Ship-stopper B — three apps are arm64-only (ISSUE-20260830-30, OPEN). Measured over ssh with a positive control in the same invocation:

XEntropy / updateRoo / AINetNode -> Mach-O 64-bit executable arm64
rdmpw3265m uname -m -> x86_64
rdmpw3275m uname -m -> x86_64

An arm64-only bundle copies to an Intel host cleanly, passes existence and checksum checks, and dies at exec with Bad CPU type in executable. Today nothing is broken — all three are installed only on the arm64 hosts. The risk is created by the deploy, not by the status quo. No universal build exists in the dev tree, so directive item 1 is a build task, not a deploy task.

replicantDB — no deploy needed. All five reachable hosts already run 1.10.1 (14) "Quiet", verified by execution (live MCP handshake returning serverInfo), not by file copy. replicantdb-ca also found the background daemon has never run under launchd on any host — the LaunchAgent invokes the app binary with --daemon, but daemon mode lives in replicantdb-cli, so it exits 0 and KeepAlive correctly declines to restart. They are fixing it; I granted the hold rather than ship 1.10.1 today and 1.10.2 tomorrow to machines that are already current.

Stale-build cleanup: deferred deliberately. A designated-requirement mismatch between hosts means a future update wipes TCC grants on the odd host out. Nothing is deleted until the install matrix is settled and each removal is matched by bundle identifier.

4. Blocked — needs Rich physically (ISSUE-20260830-25, filed by dev-76)

rdmbair15m5 is up but refuses this fleet's ssh key. Confirmed here: ICMP replies at 2.9 ms; ssh returns Permission denied (publickey,password,keyboard-interactive); control host rdmbair13m5 accepts the same key from the same session. Fault is on 15m5 — likely clobbered authorized_keys or permissions too open on ~ / ~/.ssh. Tailscale node offline, WoL did nothing, silent on the bus since 11:53. Directive item 2 (scan and consolidate 15m5 into rdmsm4x:~/dev) cannot start until it is reachable.

5. Unverified-success defect found

agy's 11:53 closeout on the bus claimed it wrote two files. Neither exists: ~/dev/_handoff/agy-master-closeout-20260830.md and ~/.agent-coordination/handovers/agy-rdmbair15m5-master-closeout-20260830.md. Also ~/dev/_handoff/staged_LogTTY/ and staged_replicantDB/ are empty directories from 00:51. rdmbair15m5 parity must be treated as UNVERIFIED, not complete.

6. Files touched

Path Change
~/dev/scripts/fleet_stability_review.zsh v1.0 → v1.1 (the fix)
~/dev/scripts/fleet_stability_review.zsh.bak-20260830-1222 new backup
~/dev/fleet/stability/stability-20260830-{1222,1223}.md validation runs
~/dev/issues/ ISSUE-20260830-28, -29 (resolved), -30

No app binary was deployed. No file was deleted. No launchd job was modified.

7. Outstanding owner actions for Rich

  1. Decide the LogTTY identifier (ISSUE-20260830-28). Recommended: keep com.eastcoastscience.LogTTY canonical — it is what is deployed and permissioned — and correct the dev tree. Cost: one Info.plist change. The alternative costs re-granting TCC on six machines for no benefit.
  2. Get to rdmbair15m5's keyboard (ISSUE-20260830-25) — ssh auth is broken and it cannot be fixed remotely.
  3. Assign a universal2 build owner for XEntropy, updateRoo, AINetNode (ISSUE-20260830-30).

No secrets, tokens or credentials appear in this record.


ADDENDUM — 12:29 EDT · two findings that arrived after the report above

A. Git push footgun — do not run a push sweep (ISSUE-20260830-31, dev-76)

Reported by dev-76, verified here independently by git config, read-only:

apps/replicantDB  branch.main.remote = roodb-legacy -> .../rdmsm4x-apps-roodb.git
apps/rooDB        branch.main.remote = origin       -> .../rdmsm4x-apps-roodb.git

Same URL. A bare git push from replicantDB writes 125 replicantDB commits into the retired rooDB repository. Both repos carry a correct backup remote that neither branch tracks. Same shape on Xentropy, scanRoo, rdLLM, rdWORKBENCH; logTTY has no push remote at all.

The correction that matters more (dev-76's, and it cuts the other way): a naive audit reads "277 unpushed commits across 26 repos" and concludes the fleet is one panic from losing everything. It isn't. Measured against the canonical backup remotes with read-only ls-remote, true exposure is ~23 commits, not 277 — replicantDB is 1 behind backup, not 125. The frightening number is an artifact of the broken upstream, and "277 unpushed after a kernel panic" is exactly the finding that triggers an emergency mass-push — the one action that fires the footgun. Measure against backup/main explicitly, never @{u}, until 31 is fixed.

B. Correction — rdmbair15m5 is NOT missing from fleet-macs.json

dev-e2 filed the failed Wake-on-LAN as a missing MAC record. That premise is wrong. Measured: ~/.agent-coordination/fleet-macs.json already holds rdmbair15m5 en0 fc:b2:14:41:58:bd — the same MAC the wake was already attempted against. arp -a shows 192.168.1.48 answering as fa:8a:00:50:6d:01, a locally-administered address (sleep proxy, or randomized Wi-Fi MAC), so the magic packet went to an address nothing listens on. No edit to fleet-macs.json is needed. Immaterial to the outcome either way: 15m5 refuses this fleet's ssh key even when awake, so it needs Rich at the keyboard regardless.

The pattern across today

Six findings, one defect class — a status that looks like a measurement and isn't:

  1. Stability reporter grading expected runs against uptime → false FAILING.
  2. Empty ack asserting acceptance on hard-blocked work, to an unreachable host.
  3. Filename and checksum checks passing over a changed CFBundleIdentifier.
  4. Existence and checksum checks passing over an arm64-only binary.
  5. @{u} reporting a 277-commit gap against the wrong repository.
  6. (dev-d5) A fixture guard asking "did anything change" against a store many agents write concurrently — it caught another agent's ticket instead of its own. On shared state the answerable question is "did anything I would have created appear".

ADDENDUM 2 — 12:33 EDT · replicantDB v1.11.0 is a real deploy candidate

replicantdb-a1 retracted its earlier "no deploy needed" — correctly, and before I acted on it. It was true at 12:22 and stopped being true at 12:30.

v1.11.0 (15) "Resident" landed on rdmsm4x at 12:31:05 EDT (commit d80626e, 424 tests, 0 failures). Fleet is on v1.10.1 (14).

Distribution-readiness checks — all pass. Build confirmed settled first (exec mtime and recent-file count stable across two samples 6 s apart), so this is not a mid-write artifact:

Check Result
lipo -archs x86_64 arm64 — universal, both Intel hosts safe
codesign --verify --strict exit 0
TeamIdentifier ZU2882L4HT, DR leaf "Apple Development: Richard Doty (S65Q255HA8)"
DR read method unanchored grep 'designated =>' — a real DR string, not the empty-output artifact
CFBundleIdentifier vs deployed com.eastcoastscience.replicantDB identical on all five hosts, all at build 14

That last row is the check that caught LogTTY, applied here and passing. No identity change, no TCC or Keychain loss.

Still not shipped. It is mechanically the first genuinely shippable artifact of the day, but it is a behaviour change, not a version bump: the fix makes a daemon that has never run on any host start real work on a ~61k-row live corpus. Recommendation to Rich is rdmsm4x first, observe an actual maintenance pass, then fan out. Go/no-go is his.

The unverified claim, carried verbatim rather than smoothed away: the verification is against a temp database. It proves the daemon starts and stays resident — alive past 18 s, daemon.lock created for the first time, second agent exits 0 leaving the first running. It does not prove the maintenance pass does useful work on production. "Daemon fixed" is not "indexing working", and nobody has verified the latter.

Fleet-wide consequence, previously unknown: the LaunchAgent invoked the app binary with --daemon, but daemon mode lives in replicantdb-cli. The app took the single-instance role, lost it to the menu-bar app, and exited 0 — which KeepAlive: SuccessfulExit=false correctly declined to retry. So the 300 s maintenance pass and FSEvents monitoring have never run under launchd on any host. Every host has had a registered daemon doing nothing. "Installed and current" was never the same as "running" — which reframes what the 11:45 directive was actually asking for, and is instance #7 of the day's defect class.


ADDENDUM 3 — 12:36 EDT · Time Machine protects ZERO fleet hosts (ISSUE-20260830-36)

dev-e2 found rdmsm4x's Time Machine stalled and asked whether the exposure was fleet-wide. I took the check. It is — and in a worse form than "they share a broken target".

Measured 12:35–12:36 EDT via tmutil, every remote call carrying a positive control in the same invocation:

Host Destination Reality
rdmsm4x Network, smb://rich@unaspro818a/timemachine stalled — 14 MB of 3.9 TB, 0.0136%
rdmpw3265m Local, "Macintosh HD" no backup at all — mount fails, Code=17 (corrected, see below)
rdmpw3275m Local, "Macintosh HD" no backup at all — mount fails, Code=17 (corrected, see below)
rdmbair13m5 No destinations configured BackupNotRunning
jdmbair13m5 No destinations configured BackupNotRunning
rdmbair15m5 unreachable ISSUE-20260830-25

Two hosts have no Time Machine at all. Two have a destination that is not a backup in any useful sense — a local destination on the same physical disk survives neither disk failure nor loss of the machine, and reports healthy forever while doing so. The one host with a genuine off-machine target is stalled. That is the day's defect class in its purest form: a backup that passes every check by being configured, and protects against nothing.

Scope limit, stated so nobody over-reads it: ~/dev is separately protected by com.eastcoastscience.devbackup and the per-repo git backup remotes. This is not "all data is unprotected". It is "a lost or failed Mac cannot be restored, only rebuilt." Worth pairing with ISSUE-20260830-31 though — the git backup remotes are themselves mis-tracked, so both safety nets are degraded at the same time.

Correction to my own contribution: I attempted a second independent byte-delta at 20 s and 25 s; my awk failed to parse the tmutil field both times and returned empty. The "0 bytes across a full 5-minute pass" figure is dev-e2's measurement, not a second confirmation. What I confirmed independently is the stalled state.

This also explains the panic, and my own blind spot. load1 of 60–100 against 71–97% idle CPU was never CPU demand — threads blocked on a stuck SMB mount count toward load average on macOS. A kernel thread wedged behind stalled network I/O is a far more plausible route to watchdogd missing check-ins for 93 s (panic 01:51:46, reboot 01:53:16) than any userspace CPU consumer. Consequence for the v1.1 fix I shipped today: a delivery grade computed from run counts reads GREEN on this host right up until it panics again, because the jobs genuinely are running. Delivery and saturation are independent signals and must not be welded into one report.

Confound flagged to the replicantDB canary

replicantdb-ca has deployed v1.11.0 to rdmsm4x only, as a deliberate canary, and is sampling an unexplained RSS climb (220 MB → 627 MB in under 3 min, not levelled off) for an hour. rdmsm4x is a compromised canary right now — wedged mds_stores and a stalled filesystem-level copy. A daemon doing FSEvents monitoring and a first index pass on a degraded I/O path may not behave representatively. An alarming curve should read as "needs a clean host to confirm", a reassuring one as weak evidence rather than a clearance. The confound cuts both ways and is now recorded.


ADDENDUM 4 — 12:40 EDT · CORRECTION to Addendum 3, and the panic mechanism

dev-e2 challenged my Time Machine framing. They were right; I re-verified both points myself before amending ISSUE-20260830-36.

My framing was wrong in the misleading direction

Measured 12:39 EDT, positive control in each invocation:

rdmpw3265m  tmutil latestbackup -> Error Domain=com.apple.backupd.ErrorDomain Code=17
                                   "Failed to mount destination."
rdmpw3275m  identical Code=17

Destination registered, ID present, mount fails. So those two hosts do not have "a backup that protects against nothing" — they have no backup at all. My wording was worse than imprecise: it hands the reader a mental model with a restore point in it, inviting the question "is the disk likely to fail?" when the correct question is "there is nothing to restore from." Withdrawn, along with the same-disk claim, which neither of us could confirm and which is moot given Code=17.

Revised count: four of five reachable hosts have ZERO machine-level recovery — rdmpw3265m, rdmpw3275m, rdmbair13m5, jdmbair13m5. rdmsm4x is the only host with any restore point.

This is instance #8, and it was mine. I checked that a destination was configured and reported as though I had checked it worked. tmutil destinationinfo is a presence check; tmutil latestbackup is the delivery check. I ran the first and reported the second.

Verified here: last successful backup 2026-08-30-012119.backup (01:21:19), newest in tmutil listbackups; boot 01:53:16; watchdog panic 01:51:46.

01:21:19   backup completes successfully
~01:21+    next backup starts and stalls on the SMB mount
01:51:46   watchdogd misses check-ins for 93 s -> panic, ~30 min into the stall
01:53:16   reboot

Stalled SMB copy → backupd/diskimagesiod/mds_stores/FPCKService pile into uninterruptible wait → load1 climbs to 60–100 while CPU is 71–97% idle (blocked threads count toward load on macOS, so this was never CPU demand) → watchdogd starves.

Two independent problems sharing a subsystem — do not merge them

dev-e2's point, and the one I'd have got wrong alone: rdmsm4x is the only host WITH a working recent backup, and it is the one that panicked. The four hosts with no backup are not exposed to this failure mode at all. The NAS stall is an availability problem on one host; the absent backups are a recoverability problem on four. Fixing the NAS does nothing for the four; configuring the four does nothing for the panic. Merging them yields one plausible remediation that half-solves both. They are separate owner actions.


ADDENDUM 5 — 12:45 EDT · Foundation Models standard verified and corrected (DOC-20260830-01)

replicantdb-ca compiled the app-baseline standard's [UNVERIFIED] Foundation Models shapes against the real SDK and routed the dev-wide library work here rather than editing it from replicantDB's context — correct routing. I independently recompiled and re-ran their probe before lifting any marker, because lifting [UNVERIFIED] on a report rather than on a compile would reproduce exactly the failure the marker exists to prevent.

Reproduced on rdmsm4x — macOS 27.0 (26A5421a), Xcode-beta MacOSX27.0.sdk, Swift 6.4. Probe committed at ~/dev/apps/replicantDB/scripts/probe_foundation_models.swift:

clean compile, no diagnostics
on-device SystemLanguageModel: AVAILABLE
PrivateCloudCompute:           AVAILABLE
contentTagging use case:       AVAILABLE
guided generation OK in 1.05s: kind=invoice confidence=95   (their run: 1.19s)
SystemLanguageModel.tokenCount(sample) = 42

The finding: a disclosure boundary hidden behind a default argument

PrivateCloudComputeLanguageModel is public, publicly constructible via init(), conforms to LanguageModel, and reports AVAILABLE. LanguageModelSession accepts model: some LanguageModel. The difference between on-device inference and sending a user's file contents to Apple's servers is one argument, at one call site — no entitlement gate, nothing in the type system, and the safe and unsafe calls look identical at the point of use and in review.

For replicantDB, DataRoo, ReceiptRoo and bookmarkROO this is a disclosure boundary, not a performance tradeoff. Off-device inference may be right for some feature one day; it is never right by accident, and it is Rich's decision, not a default argument's.

Two API-shape corrections — the standard was wrong; code written against it would not compile

  1. tokenCount(for:) is on SystemLanguageModel (macOS 26.4+), not LanguageModelSession.
  2. No static context-window constant exists. The standard's guessed number was deleted, not corrected — measure with tokenCount and handle the overflow error, which carries contextSize and tokenCount so you can report exactly what did not fit.

Changes to ~/dev/lib/app-baseline/PLATFORM_INTELLIGENCE_AND_PERFORMANCE.md

Backup .bak-20260830-1244. Markers lifted on what was compiled. New BINDING section "The disclosure boundary", placed before "Where it genuinely helps", showing the safe and unsafe call sites side by side. Three binding rules: bind SystemLanguageModel.default explicitly at every session construction · ship a test that fails if a PCC model can reach a session · off-device inference is an explicit, reviewed, user-visible decision. New section "Capacity" recording 1.05–1.19 s/item ≈ 20.7 hours single-threaded over 62,706 items, so nobody scopes on-device classification as a corpus-wide sweep.

Deliberately still unverified: "MLX has no ANE target" elsewhere in the standard was not probed and must not be upgraded on the strength of this work. The probing session declined to confirm what it had not measured while everything around it was being confirmed — the most copyable behaviour of the day.

Open, needs an owner: audit every app that constructs a LanguageModelSession for an explicitly bound model (none audited yet — the standard now requires it, nothing enforces it), and write the failing test as a shared fixture in ~/dev/lib so apps inherit rather than reimplement it.

Instance #9, and the subtlest: an API that exists, constructs, and reports AVAILABLE — where availability was never the question anyone should have been asking.


ADDENDUM 6 — 14:09 EDT · rdmbair15m5 returned. Directive item 2 CLOSED, and the backup picture completed.

The host came back — and neither of our diagnoses was right

rdmbair15m5 became reachable at ~14:05 EDT. dev-76 then pulled the timeline that settles it:

12:01:32  boot
12:08–12:29  SSH DENIED from all four hosts
12:49     richh logs in AT THE CONSOLE
14:05     SSH works — authorized_keys mtime 2026-08-21, UNCHANGED, perms correct, our key present

On a FileVault Mac, sshd cannot read a user's authorized_keys until that user has logged in at the console since boot. It still offers publickey and still returns Permission denied (publickey,password,keyboard-interactive) — indistinguishable from a bad key. Both of us diagnosed clobbered keys or bad permissions; all four particulars were false, and the remediation would have changed nothing while SSH came back anyway — which would have "confirmed" the wrong cause. FLEET.md §1 already documented this failure mode and we both reached past it for a more interesting hypothesis. The lesson is check the documented failure mode before diagnosing a novel one, not "we misread the error". No action on key material, ever, for this; sudo fdesetup authrestart is the planned-reboot answer.

Directive item 2 — satisfied by measurement, not by consolidating (TASK-20260830-01, resolved)

Read-only assessment, 14:06–14:09. Nothing pulled, moved, promoted or deleted. rdmbair15m5:~/dev = 18 GB, 191,759 files. It holds no unique work.

Exactly two repos carried anything not on a remote, and both counts were misleading — resolved by direct SHA comparison rather than trusted:

Repo Raw count said Truth
net/cloudflare-access-passkeys 6 unpushed All 6 SHAs present on rdmsm4x — identical history
apps/LogTTY 1 unpushed Stale build-18 snapshot vs hub's 111 commits at build 21 — do not merge

~/dev/rdmbair15m5 — the directive's second path — is an empty directory. Half the target did not exist.

Why the counts weren't trusted: git log --branches --not --remotes compares against whatever remotes are configured, and ISSUE-20260830-31 established those are mis-tracked. A count from it can be right about the number and wrong about the meaning. Both were.

This supersedes agy's 11:53 closeout, which asserted parity complete and cited two files that do not exist. Its conclusion was right; its evidence was not.

The backup picture, completed — and a new single point of loss

rdmbair15m5  tmutil latestbackup -> "Failed to find any backups found for current machine"

Configured against the correct NAS, never completed a backup. Final count, all six measured: five of six hosts have ZERO machine-level recovery. rdmsm4x is the only host with any restore point, and its current backup is the stalled one that preceded the panic.

Compounding: net/cloudflare-access-passkeys — 6 commits, identical on rdmsm4x and rdmbair15m5, no git remote on either host, and no working Time Machine on either. That project exists on two local disks and nowhere else. It is not covered by ISSUE-20260830-31, which concerns repos with mis-tracked remotes; this one has none at all.

FLEET.md §1 — measured hardware inventory added

Backup .bak-20260830-1406. RAM, disk and macOS build for all six hosts, measured. Surfaced two things: jdmbair13m5 has 135 Gi free of 926 Gi — an order of magnitude less headroom than any other host, so nothing should be promoted there — and the table's macOS column for it was stale (said 26.6.2, actually 27.0). The two Intel hosts are on 26.7 while all four Apple Silicon hosts are on 27.0.


ADDENDUM 7 — 20:44 EDT · Rich transferred LogTTY ownership. A shipping trap, and a rule conflict.

Rich, 20:38 EDT: "resume all - make claude@rdmsm4x take ownership of LogTTY if that is the question - approved for all". Taken.

🛑 dist/logTTY.app is a trap — broadcast as do-not-ship (bus 20260830-204140-4F31D7BF)

It was rebuilt at 14:51 today, after this morning's measurement, and is a regression in every dimension at once:

This morning (12:25) Now (14:51 rebuild)
Build 20 18 — fleet runs 19, HEAD intends 21 → downgrade
Arch universal x86_64 arm64 arm64 only → dies on both Intel hosts
codesign --verify --strict clean "a sealed resource is missing or invalid"
Signature Team ZU2882L4HT ad-hoc, no team → TCC re-prompts every launch

Root cause: BUILD_NUMBER is reverted 21 → 18 as an uncommitted working-tree change (HEAD 23f9a89 is literally "bump BUILD_NUMBER to 21"), and the build was packaged from that dirty tree. Same class as the 08-29 23:09 sweep that left replicantDB unsigned.

Checked before alarming anyone: the repo is INTACT — 111 commits, HEAD 23f9a89, no reset in the reflog, and 3d3c908 is a legitimate ancestor rather than spoke content pulled backwards. Known-good reference is /Applications/LogTTY.app on rdmsm4x: build 19, universal, strict-clean, ZU2882L4HT. Not dist/.

Identifier decision — settled by the house rule, not by my judgement

~/dev/apps/CLAUDE.md §5: "Bundle IDs: com.eastcoastscience.<Product>." Surveying where the identifier is set rather than mentioned, the repo is already overwhelmingly capital-L — build_and_run.sh (which packages the app), verify_product_identity.sh, deploy_release_remote.sh, the packaging test, the CloudKit string, the mobile target. Lowercase survives in one subsystem (the privileged helper).

Decision: com.eastcoastscience.LogTTY. The product name stays "logTTY" — the rename was a product rename and identifiers are opaque plumbing users never see, so keeping capital-L costs it nothing. No host loses TCC or Keychain. The helper subsystem reconciles to it.

My own correction, twice over: I first recommended capital-L, then withdrew it on finding the rename was deliberate, then restored it on finding where the identifier is actually set. And I told a peer I could not find the CloudKit rule in apps/CLAUDE.md — it is there, in §5; I searched §3 and concluded from its absence. That is the day's own defect class, committed by me, for the second time.

PLATFORM_TARGETS.md — a real conflict, reconciled rather than overwritten

Rich set a new standing rule (via replicantdb-a1): universal2 is a build gate, lipo -archs fails the build. But PLATFORM_TARGETS.md already carried Rich's 2026-08-22 decision that Intel ships as "a separate scheme with its own deployment target, not a universal binary". A blanket universal mandate reads as overturning that, and silently overwriting one recorded decision of Rich's with a later relayed one was not something to do quietly.

The two rules answer different questions, which is what dissolves it:

Question Answered by
Does the binary contain code this CPU can execute? the lipo gate
Will this OS version load it at all? the 26.7 scheme

Intel support requires BOTH — an app with a 27.0 deployment target cannot run on Intel even with an x86_64 slice, because no Intel Mac will ever run 27.x. So lipo showing two slices is not evidence of Intel support. Both rules stand; neither is overturned. Written into the file as an explicit "read both" table. Backup .bak-20260830-2043. Host table also refreshed — builds had moved and jdmbair13m5 was missing entirely.


ADDENDUM 8 — 20:52 EDT · CloudKit: the real blocker, named (SEC-20260830-04)

Rich, 20:47: "I've been asking for cloudkit for weeks." Root cause found. It is not what apps/CLAUDE.md §3 has been telling every session for weeks.

Measured — every provisioning profile on rdmsm4x, decoded with security cms -D

Profile App ID Container Platform
net.dataroo.RTTy.mobile ZU2882L4HT.net.dataroo.RTTy.mobile iCloud.net.dataroo.RTTy iOS
com.eastcoastscience.Tyrell ZU2882L4HT.com.eastcoastscience.Tyrell iCloud.com.eastcoastscience.Tyrell iOS
team wildcard ZU2882L4HT.* none iOS

Three profiles, valid into 2027. All three iOS. There is not one macOS provisioning profile on this machine, and no CloudKit container exists for any macOS app. Every app on this fleet is a Mac app. The transports have been ready for weeks — 440 + 73 + 218 lines — pointing at containers that were never created.

The trap: the team wildcard ZU2882L4HT.* carries no containers and structurally cannot — Apple forbids iCloud containers on a wildcard App ID. So signing succeeds against it, builds pass, and CloudKit silently does nothing.

Why every attempt blamed something different: the layers fail in sequence — no container → no capability on the App ID → no macOS profile → no entitlement in the app — and each attempt hit whichever layer it reached first and reported that. §3's "the blocker has always been schema promotion and signing" is half right: it is signing, specifically macOS App IDs and profiles that were never created. Nothing can be promoted to a schema in a container that does not exist. That sentence has been routing sessions to the wrong layer; corrected in apps/CLAUDE.md §3 (backup .bak-20260830-2051).

Done tonight — everything not needing Rich's Apple ID

Needs Rich — ~20 minutes at developer.apple.com, LogTTY first to prove the flow

App ID (macOS, explicit, iCloud capability) → create container iCloud.com.eastcoastscience.LogTTY typed once, exactly, permanent → assign → macOS App Development profile → install. Or let Xcode's Automatically manage signing do it once the container exists. Then replicantDB (its entitlements are already correct and waiting) and RTTy (adopt a new container; the dead iCloud.net.dataroo.RTTy cannot be removed).

Verification rule recorded: never report CloudKit working until a record written on one Mac is read back on another. Every layer passes its own check while the layer below is missing — instance #10 of the day's defect class, and the one that cost the most calendar time.

Also this pass

~/.claude/CLAUDE.md — platform targets added to the global macOS rules at Rich's request: deployment targets, the two Intel hosts confirmed by Rich on 26.7 (25G224), the universal2 build gate, the necessary-but-not-sufficient trap, and the bundle-ID/CloudKit/signing rules. Backup .bak-20260830-2048.


ADDENDUM 9 — 22:04 EDT · I got one wrong. rdmbair15m5 consolidation REOPENED (TASK-20260830-06)

replicantdb-a1 routed an intake item that contradicted a conclusion I had published. It was right and I was wrong. Corrected in the ticket rather than quietly.

What I wrote (TASK-20260830-01): "rdmbair15m5 holds NO unique work. Parity is now VERIFIED by measurement." What I actually measured: no git repo under rdmbair15m5:~/dev carries commits absent from the hub. That narrower claim still stands. The gap: I generalised a repo-scoped instrument into a claim about all work. My scan walked find ~/dev -name .git -maxdepth 4, so a non-git directory was invisible to it by construction, and ~/Archived-migrated-2026-08-13 was never in scope at all — it is not under ~/dev. Both exclusions contained real content.

Measured 22:02–22:03, positive control per call

Path State
~/Archived-migrated-2026-08-13/ 89 GB, 605,011 files — never triaged by any pass
~/dev/obsidian_fleet_sync not a git repo — 234 files; 190 tests; its own doc names a canonical home ~/dev/agy/obsidian_fleet_sync/ that does not exist
~/dev/fleet-agent-registry not a git repo — 15 files, 92 MB
ticket store 378 tickets vs this host's ~113 — needs the ticket CLI, not a file merge

a1 reported the archive as 48 GB; measured 89 GB. Correcting upward — the finding is theirs.

Relay manifest verified independently: HANDOFF.md, 7,707 bytes, shasum -a 256 = b7bce9044fd9f0921423f4000808969089d8b2d54ff35665a5572cbd60994f1d — exact match to what codex@rdmbair15m5 published on both ends. Recorded RECEIVED, not accepted: checksum parity proves the file crossed intact and nothing about whether its content is correct or complete.

Sequence — deliberately not starting with the 89 GB

obsidian_fleet_sync first (smallest, highest confidence, closes a doc that already claims a home that does not exist) → fleet-agent-registry but identify the binaries first: 92 MB in 15 files usually means a database or cache, which should not become canonical → Desktop HTMLs, almost certainly regenerable exports → ticket store, CLI only → the archive LAST and by sampling. Establish what fraction is unique before any transfer; it predates the canonical-root discipline and is likely mostly duplicates. And nothing from it should land anywhere reaching jdmbair13m5, which has 135 Gi free of 926 Gi.

Instance #11, and this one is mine

Same shape as the error I flagged in someone else's Time Machine premise four hours earlier, and as my own tmutil destinationinfo report: I confirmed the narrow technical fact and reported the broad operational conclusion. The rule now recorded fleet-wide — "a search that matches nothing is not evidence of absence; confirm the search could have found it" — applies equally to a search that matches everything it can see. A repo scan cannot see a non-repo. State the scope of the instrument in the conclusion, every time.


ADDENDUM 10 — 22:09 EDT · Item (a) closed by verification, not by promotion

I was one command from creating a third copy of a project that was already fully promoted.

Measured: ~/dev/apps/ObsidianFleetSync is a git repo, 2 commits, 241 files — and the file-set difference against rdmbair15m5:~/dev/obsidian_fleet_sync is zero. Not zero-excluding-scratch; zero, including every .agents/ working file. The hub copy is a strict superset of the spoke copy. Promotion was completed 2026-08-29 under ISSUE-20260829-04, and that ticket was accurate.

Why it convincingly looked unpromoted — four stale artifacts all pointing the same wrong way

  1. The project's own docs declared "Canonical home: rdmsm4x:~/dev/agy/obsidian_fleet_sync/" — a path that does not exist. It went to apps/, not agy/.
  2. ~/dev/todo/items/20260826-obsidianfleetsync-never-promoted-no-git.md still said never promoted, no git — superseded three days later and never closed.
  3. A partial duplicate at ~/dev/apps/obsidian_fleet_sync (lowercase — a genuinely different inode, 342109867 vs 340193522, not a case-insensitivity artifact): 72 files, not a git repo, and a strict subset. Anyone checking that path finds a non-repo and concludes "not promoted". That is almost certainly what happened.
  4. A fourth copy staged at ~/dev/_inbox/rdmbair15m5/obsidian_fleet_sync.

Every link in that chain is individually true and the conclusion is false. The check that broke it was comparing file sets between hub and spoke rather than checking whether a named path existed. Existence of a path is a presence check; set difference is a delivery check. Instance #12 — and the first one caught before acting rather than after.

Fixed at the source (commit 920dc4e)

The stale canonical-home claim is corrected across SESSION-STATE.md, ISSUES.md, RECONCILIATION-20260823.md and OBSIDIAN-CONSOLIDATION-AND-ROODB-PLAN-20260823.md — because it was actively harmful, not untidy: it caused one false "unpromoted work" finding tonight and would have caused more. The superseded todo item is closed with the evidence.

(My first substitution ate a path separator and produced ~/dev/apps/ObsidianFleetSyncRECONCILIATION-20260823.md. Caught on the verify pass, restored from backup, redone correctly. Noting it because the verify pass is the only reason it did not ship.)

Nothing deleted. The two redundant copies are provably lossless to remove and are recorded for removal — but a directory deletion at 22:09 unattended buys nothing when the finding is already recorded. And explicitly: do NOT create ~/dev/agy/obsidian_fleet_sync/. That is the trap.

Credit where it is due

replicantdb-a1 raised this in good faith from a read-only scan and the signal genuinely looked like unpromoted work. The finding was wrong; raising it was right — it surfaced four stale artifacts that were actively misleading every agent that looked, and those are now fixed.


ADDENDUM 11 — 22:12 EDT · Withdrawing my own jdmbair13m5 constraint (Rich reviewed it)

dev-76 walked back an operational rule they had relayed to Rich. The inference underneath it was mine, and it is now withdrawn in both places I put it.

What I wrote in FLEET.md §1: jdmbair13m5's 135 Gi free is "the only one where a large sync or index could plausibly run it out." And in TASK-20260830-06 step (e): "nothing from this archive should be promoted anywhere that lands on that host."

Both withdrawn. Rich reviewed it and 134 Gi is fine for that host's role. Two reasons, both sound:

  1. jdmbair13m5 is a spoke/scratch node. Canonical is rdmsm4x by definition — so "don't promote canonical there" warned against something the architecture already excludes.
  2. 86% used on APFS is not the 86% the HFS+-era instinct reacts to. APFS does not fragment-degrade near capacity, and local snapshots purge under pressure.

What survives, and is worth keeping: jdmbair13m5 is the fleet's only 1 TB host (the others are 4 TB or 8 TB) and the only one above 50% used — 1.0 TB APPLE SSD AP1024Z, 767 Gi used / 134 Gi available, no external disks. FLEET.md now records that as "inventory, not a constraint" and marks the prohibition withdrawn in the file rather than deleting the sentence, so a later reader cannot re-derive it from a bare number.

A genuinely different route to the same failure

Every other instance today was one agent over-reading its own instrument — my repo scan, my tmutil destinationinfo, the anchored designated-requirement grep. This one only exists because the claim crossed between agents.

I attached a hedged inference to a measured number, in a file five other agents read as ground truth. dev-76 relayed it to Rich as a firm rule without measuring it. Rich challenged it — "4TB drive - you sure it's not your script?" — and only then did anyone verify. The number held; the conclusion never had support.

A verified number does not confer verification on the inference travelling beside it. It arrives as context attached to something that was checked, so it never presents as a claim needing checking. That is why neither of us checked it.

One concrete fix on my side: the 4 TB Rich was thinking of is rdmbair15m5, which sat two lines above jdmbair13m5 in my own report. Adjacent rows with similar names and very different numbers invite misreads — worth guarding against in how per-host tables are laid out, not just in what they say.