Fleet dev_update sweep: no real runs anywhere 09-09→09-21, fixed — and an incident my own verification caused
When: 2026-09-21 15:34:44 → 2026-09-21 17:22:49 EDT
(both read from the clock) · Host doing the work:
rdmsm4x · Scope: all six fleet Macs
Tickets: ISSUE-20260921-07 (this work) ·
ISSUE-20260920-04 (re-scoped, open) Commits
(~/dev/scripts): 5660583 713f8d8
617a1c4 ce3e1c6 d0fbf22 (+
SESSION-STATE d74f64c 6bd6917
4e3c357)
One line: the nightly fleet update had silently stopped doing real updates — no spoke after 2026-09-05, the hub after 09-08; it now runs again (5 of 6 hosts verified today) — and while proving that, my verification run briefly switched the hub's Xcode under another agent's App Store build, which is now impossible.
What was wrong
fleet_dev_update.zsh (launchd 03:07 on rdmsm4x) gives a
host a real run only if its dry run passes. v1.0 required exit
0 — written for the bare payload's codes (0 ok / 1 issues / 3
concurrent). On 2026-09-04 ~/dev_update.zsh became the
self-contained wrapper, whose contract is 0 clean / 10 warnings / 20
needs a person / 30 critical. Every spoke carries findings that are
permanent until Rich acts (Developer ID .p12, Full Disk Access toggles,
Xcode sign-in), and every Intel dry run raised a warning by
construction, so exit 0 was unreachable. No spoke got a real run
after 2026-09-05 and the hub none after 09-08; from 09-09 to 09-21 all
78 host-nights read "real run gated off". Two dry runs also
never finished in the 600s cap: rdmbair15m5 hung in an unbounded
docker info (8 of 10 nights), and rdmsm4x ran out of time
in npm doctor … cache, which also garbage-collects the npm
cache — a write, inside a dry run.
What changed
- fleet_dev_update.zsh 1.1 — gate reads the wrapper contract: real run on 0/10/20, skip on 30 (host unfit), timeouts, ssh failures, anything unknown. Unit-tested for every code. Dry-run cap 1200s as a backstop. The report labels each code and shows the wrapper's one-line verdict per host.
- dev_update v2.01 / payload 2.15 (new file; v2.00
kept as the rollback), on all six hosts, payload sha
f87802488dca:- Docker: every call bounded (20s probes, 600s prune); a silent daemon is info, not a warning.
- npm doctor: no
cachecheck in a dry run; bounded at 300s; a timeout is "not judged", never "passed". ollama servegets its own log (it had held a 2026-09-18 dev_update log open for three days).timeout -keverywhere (a TERM-ignoring child is KILLed); exit 137 reported as a timeout.--no-system, like--no-mas, is the operator's flag → a warning, and a pending beta is named as one (the hub's only "update" was macOS 27.2 Beta 2 and the old text saidsudo softwareupdate -ia).mas outdated: only real "( -> )" lines are apps (Spotlight notices had become a 5 KB app list); MAS_NO_AUTO_INDEX=1in a dry run so mas doesn't start Spotlight indexing.- "dry run: nothing rebuilt" is info, not a warning.
- Xcode selection (after the incident below): the
orchestrator's
rmfix_xcode_selectis report-only; the payload's ownsudo xcode-select -sruns only when attended AND no~/dev/_handoff/*/XCODE-SWAP-ACTIVEmarker AND the target isn't named*swapped*/*parked*AND no xcodebuild is running.
Incident — caused by this session, contained, fixed
At 16:24 I ran the sweep by hand to verify the gate.
Its rdmsm4x real run's remediation ran
sudo -n xcode-select -s onto the parked
/Applications/Xcode-27.0-swapped-for-rtty-514.app (~16:26)
— while the RTTy lane (session c13f617a) was building RTTy
0.3.839 Build 514 for the App Store with Xcode 26.6
deliberately swapped into /Applications/Xcode.app
(DEC-20260903-07). Restored at 16:27:29 and the lane
told. Its confirmation: no harm — that construct failed closed on an
unrelated validator check before any export existed; the re-run built
under 26.6 (DTXcodeBuild 17F113 / macosx26.5) and Build 514 is
VALID in App Store Connect. The guard above was verified live
while the lane's swap marker was present ("left alone").
Verification (numbers, not "green")
| check | result |
|---|---|
sweep --dry-run-only
16:16 |
6 of 6 dry runs completed; 6 of 6 fit to update (v1.0: 0 of 6) |
| full sweep 16:24:41 → 17:20 (3341 s) | 5 of 6 real runs completed; rdmbair13m5 timed out at the 1800s cap |
| rdmsm4x dry run | 24–31 s, exit 10 (at 03:07 it was killed at 600 s) |
| rdmbair15m5 | real run completed — its dry run had timed out on 8 of the previous 10 nights |
--verify-embedded |
OK on all six |
Still open
- ISSUE-20260920-04 (rdmbair13m5): Homebrew 7's
sandboxed operations hang there —
postinstallyesterday,extracttoday (two unrelated brews stuck onsandbox_operation.rb extract, one of them an interactivebrew upgrade --greedyin a Terminal window). That host can't complete brew upgrades until it's diagnosed. Related: rdmpw3275m'sherdr→llvm@21pour failed inside Homebrew's sandbox (Tier 3); herdr is pinned on rdmsm4x but not on the Mac Pros. - Rich only: Developer ID .p12 on four spokes; Full Disk Access for claude/codex/agy on rdmpw3275m and jdmbair13m5; Xcode Apple ID + first launch on jdmbair13m5.
- Declined by Rich (~16:05): an agent-run
sudo mv/chownin the Homebrew prefix on rdmbair13m5. Not retried.
Undo
ln -sfn dev_update_v2.00.zsh ~/dev/scripts/dev_update.zsh
on any host; git revert the fleet_dev_update.zsh commits
(v1.0's gate would again skip every real run).