Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260909-2225-rtty-build-505-developer-id-notarized-six-hosts

rdmsm4x — RTTy Build 505: first Developer ID, notarized, stapled build on all six Macs — 2026-09-09

One line: RTTy Direct 0.3.830 build 505 now runs on all six fleet Macs signed with Developer ID, notarized by Apple and stapled — and it took an app fix to get there, because the first Developer ID candidate launched, collected data, and drew no window at all.

Scope

Hosts changed: rdmbair15m5, rdmbair13m5, jdmbair13m5, rdmpw3265m, rdmpw3275m, rdmsm4x — all six on /Applications/RTTy.app 0.3.830 build 505, Developer ID, notarized, stapled.

What changed and why

Rich's decision DEC-20260909-03: every fleet app ships Developer ID + hardened runtime + --timestamp, notarized and stapled, TCC re-prompts accepted. For RTTy that needed three PRs and one Apple portal action by Rich.

  1. PR #40 — the release gate now judges notarization, not the certificate's name: the signing mode is derived from the profile, stapler validate must pass, spctl must report source=Notarized Developer ID, get-task-allow must be absent (notarization refuses it), and the signed CloudKit environment must equal what the profile authorizes.
  2. PR #41 — resolve_codesign_identity.sh resolves by certificate KIND as well as by hash. Widening the match would have broken a deliberate guard forbidding a Developer ID certificate from satisfying a request for a development identity, so the kind is an explicit argument that defaults to apple-development.
  3. PR #43 — the fix that made shipping possible (DEC-20260909-05), below.
  4. Rich created the Developer ID provisioning profile RTTy Direct Developer ID (a2b99f8b-8139-4d5a-86af-8bc866b87719) carrying the iCloud container; without it a Developer ID build could not launch at all.

The failure that mattered, and how it was found

The first Developer ID candidate ran and probed — four ping children — but had no UI: 0 menu bars and 0 windows after 75 seconds, against 2 and 1 for the installed Build 504 minutes later.

sample named the line, twice:

ObservationStore.start() -> probeAllTargets -> publishCloudProjectionIfAvailable
  -> currentFleetSnapshot -> StableDeviceIdentityStore.deviceID()
    -> SecItemCopyMatching -> securityd decrypt     <- parked here

A synchronous Keychain read on the main actor. Under the new signing identity, the device-identity item's ACL still named the old one, so the read waited for an authorization that never came and applicationDidFinishLaunching never completed. CloudKit was downstream and never reached — the first diagnosis (an unpromoted CloudKit schema) was wrong, and the stack said so.

The fix has two halves, because either alone leaves the trap armed: the identity is read through loadDeviceIdentityOffMainActor (a helper that already existed for this exact hazard and that this one call site did not use), and the fleet projection publish left the probe path entirely — single-flight, a real 10 s timeout raced against the attempt, 30 s→15 min exponential backoff, failures recorded in the CloudKit state surface that already existed and logged once per distinct reason.

Two-direction proof, same identity and profile: before — 0 menu bars, 0 windows, no latency value, still 0 after 75 s. After — 2 menu bars, 1 window, "Overall average latency 18 ms" within 20 s.

Artifact

Fact Value
Version / build 0.3.830 / 505, commit cf15adc947d7518533305715b6ba38cd47caca9b
Executable SHA-256 00dcb4b7d91affc55a54a8974235692e7c4cffbdc130cf1b5f48e4346a931af4
Archive RTTy-direct-0.3.830-build505-cf15adc.zip, 13,361,882 bytes, a7323ef7…7ddf
Signing Developer ID Application: east coast science, llc (ZU2882L4HT), hardened runtime, secure timestamp
Profile a2b99f8b-8139-4d5a-86af-8bc866b87719, SHA-256 0efc0851…1f02
Notarization submission 28d9d6eb-98e8-414e-a583-983f3658f1d3, Accepted, stapled
Architectures x86_64 arm64

Gates and tests

Ten of ten gates PASS, rc=0, 78 assertions. Native and Rosetta each 1058 of 1063 discovered XCTest cases, 0 failures. Gate 7 — the live menu-bar latency probe that caught the hang — passes. New unit tests (XCTest, so the gate actually runs them): a never-returning writer is abandoned inside the timeout; a CloudKit error is recorded as unavailable with its reason; a second schedule does not start a concurrent attempt; a failed attempt backs off and later resumes; the backoff grows and is capped; the device identity is read off the main actor.

Deployment and verification

Canary rdmbair15m5 first, hub last, 22:09–22:17 EDT, Build 504 (5016c2b8…) armed as rollback on every host at /Applications/.RTTy.app.rollback-build504-cf15adc.

Verified on all six by execution: 0.3.830/505, x86_64 arm64, exe 00dcb4b7…, Authority Developer ID Application, stapler validate ok, spctl → source=Notarized Developer ID, profile 0efc0851…, one RTTy process each running from /Applications/RTTy.app. Live latency confirmed under the new identity on the hub (15 ms) and the canary (18 ms).

Two installer defects found — fix before the next deployment

Both are in ~/dev/_handoff/rtty-build50*/release/*/tools/deploy_build50*_fleet.zsh, not in the repository:

  1. trap 'status=$?; …' EXIT assigns to status, which zsh makes read-only — the receipt shows zsh: read-only variable: status. The rollback-on-failure path therefore never runs. On a real failure this installer leaves the new build in place instead of restoring the old one. Rename the variable (e.g. exit_status).
  2. The 90-second acceptance window is too tight for a first launch under a changed code identity. rdmpw3275m returned rc=1 on it while the host was in fact correct and collecting: measured directly afterwards, Build 505 installed, 4 ping children, and statistics.interval_end advancing 1789006637.9214 → 1789006697.99218 over 45 s.

Backups and how to undo

Every host: sudo mv /Applications/.RTTy.app.rollback-build504-cf15adc back over /Applications/RTTy.app after quitting RTTy, then open it. Rollbacks for Builds 501–503 remain where they were; nothing was deleted.

Repository: main at 80e5a05, tag v0.3.830-build505 at cf15adc, pushed to origin, backup and fleet. Evidence: docs/release-evidence/build-505/fleet-deployment-2026-09-09.md.

Outstanding owner actions

  1. Promote the CloudKit Production schema. A Developer ID profile authorizes only the Production environment, and RTTy's Production schema has never been promoted, so the fleet projection is expected to report unavailable and back off. That is now bounded and logged rather than fatal — but it is not a working projection.
  2. TCC re-prompts are expected once per host because the code identity changed. Every host launched and collected, so nothing blocking was observed, but privacy prompts may appear at the console.
  3. App Store / TestFlight remain blocked on the Apple Distribution certificate and App Store profiles.

Note: Build 505 is commit cf15adc and does not include PR #42 (the application-flows redesign), which merged to main afterwards.

No secrets are recorded here; keys and profiles are referenced by name, UUID and location only.