dev_update v2.00 / payload 2.14 — self-heal the recurring findings, and the TCC read that silently stopped a host
When: 2026-09-20 20:19 → 21:30 EDT · Host
doing the work: rdmsm4x · Scope: all six fleet
Macs Ticket: ISSUE-20260920-03 ·
Commits: 91bc510, 8b6c283,
a28c277 in ~/dev/scripts
One line: every nightly dev_update run on this fleet was
re-reporting the same handful of findings because nothing in the script
could clear their cause — and while verifying that,
rdmbair13m5 turned out to have silently stopped updating
two days earlier.
What changed
~/dev/scripts/dev_update_v2.00.zsh (new file;
dev_update.zsh symlink repointed from
dev_update_v1.99.zsh). It carries
dev_update.zsh payload 2.14 in its
sha-pinned heredoc. --verify-embedded passes on all six
hosts.
The internal v199 identifiers (log filename, self-pin
temp path, payload cache dir, footlights state key) were deliberately
not renamed: they key state, and renaming them
mid-rollout would orphan it on half the fleet for no gain.
1. Zombie cask receipts — self-healed, once per cask
brew upgrade --cask aborts a cask whose staged
.app was deleted by hand
(It seems the App source '<path>' is not there),
purges the version it just fetched, and keeps the stale receipt — so
brew outdated --cask lists it again next run and the
identical error is raised for ever. Three hosts at once:
backblaze-restore on rdmsm4x (its Caskroom version dir an
empty 0B directory), taskexplorer on both Mac Pros.
brew_cask_selfheal() reinstalls the named casks with the
reversible brew reinstall --cask --force. Done once
per cask — a cask that disappears again after a successful
repair was removed on purpose, and re-downloading it nightly would be
the script overriding the operator; the second occurrence names both
choices instead.
2. The cask batch is judged by a re-check, not by brew's exit code
brew upgrade --cask exits non-zero for a per-cask
failure even when the upgrade landed. On jdmbair13m5 the
microsoft-edge post-install chmod -R a+rX,go-w
was denied by macOS App Management, brew called that a failure and its
own rollback failed too — and the measured end state was the new app in
/Applications, the new receipt in the Caskroom, and nothing
outdated. The batch now re-reads brew outdated --cask:
nothing outdated means the goal was met. When something IS still
outdated, an Operation not permitted in the output is named
as macOS App Management (a GUI consent toggle no script will grant)
rather than "exit 1, see the log".
3. brew doctor — the mechanical half, in the payload
Three classes Homebrew names its own one-command reversible fix for
are applied before the verdict, and
brew doctor is re-read so the verdict describes the result:
unlinked kegs (brew link), invalid Caskroom metadata
(brew reinstall --cask --force), and broken symlinks in the
prefix (moved to an archive, never deleted).
Never brew link --overwrite: the blocking file is
usually a hand-installed gem and --overwrite deletes it. A
blocked link reports Homebrew's own command instead. Note
brew link --dry-run does not predict this
— it returns 0 and "Would link" while the real link fails on an existing
file.
The keg list is read from the line after
Run `brew link` on these:. The orchestrator's existing
rmfix_brew_doctor terminated on the first blank line after
the Warning, but that stanza carries a blank line before the
heading, so it collected nothing — which is why ruby stayed
unlinked on every host for nine days with a fixer for it already in the
tree. Fixed in both places.
4. Intel support-tier notice
Homebrew added
Warning: You are using macOS on Intel x86_64 in Sept 2026 —
the same class as the macOS Tier-2 notice already filtered. The old
filter matched only the macOS-version form, so both Mac Pros reported
"brew doctor found issues" every run whose entire content was Homebrew
saying Apple dropped Intel. The filter now matches the
Warning: header of both; matching the header and not the
prose keeps it robust against rewording.
5. Severities corrected to what an unattended run can act on
uv cache pruneis skipped, naming the holder, when a liveuv runowns the cache lock. The Claude Desktopmacos-mcpextension holds it from login to logout (pid 52814 on rdmpw3275m,lsof-confirmed), so the prune waited out the 300s lock timeout and exited 2 on every run. Otherwise bounded byUV_LOCK_TIMEOUT=60.--no-mas was setis a warning, not an issue:fleet_dev_update.zshpasses--no-masmandatorily on every SSH run (documented:mas upgradehangs 30+ min on a headless sudo prompt). Reporting the operator's own setting as Action Required is a nightly false positive by construction.- pip — packages in an interpreter this user cannot
write are a warning with the real remedy; and a package another
installed distribution pins exactly
(
pydantic 2.13.5requirespydantic_core==2.46.5) is a constraint, not an action item. Classified offline from installed metadata.
6. The serious one: a system-TCC read with no deadline, holding the run lock
rdmbair13m5 had silently not updated since 2026-09-19
03:13. A dry run from that morning was still alive 41 hours later, 0%
CPU, blocked on
sqlite3 /Library/Application Support/com.apple.TCC/TCC.db 'select ... from access'
with TCC.db never even opened. Reading the
system TCC db needs Full Disk Access, and without it
the read does not fail — it blocks on a consent
decision nobody answers in an ssh or unattended session.
_perm_sql had no deadline on either the local read or the
ssh-localhost fallback.
That run held the single-run lock, whose gate only asked
kill -0. A wedged holder is alive for ever, so every later
run was refused with "another fleet_update is running" — and the run
that reports it is the run being refused, so the only symptom
was a log date that quietly stopped advancing.
_fu_timed()bounds both reads at 20s (FU_SQL_TIMEOUT), with a shell watchdog for a host without coreutilstimeout— the protection must not itself depend on an optional formula.- The lock gate group-kills a holder only when both
hold: the lock file is older than 12h (
FU_LOCK_WEDGE_S; the Intel tuning phase has measured 5h40m, so the bar is well clear of a real run) and the holder's command line is one of ours. Group kill, because the wedge is the child — killing the parent alone orphans the blockedsqlite3still holding on.
7. An unattended cask upgrade gets no stdin
A cask with a pkg artifact shells out to sudo, and one
that has to close a running app asks first. With a terminal on stdin
that prompt just sits there. Measured on rdmbair13m5 21:19-21:35:
brew upgrade --cask --greedy-latest at 0% CPU for 15
minutes with fd 0 on /dev/ttys013. When stdin is not a
terminal the step now runs </dev/null, so a prompting
cask fails — bounded, named by the re-check, reported —
instead of stalling the run. An attended run keeps its terminal.
-t 0 and not ~/.agent-coordination/UNATTENDED:
that marker exists on rdmsm4x only, and permanently, so keying off it
would strip prompts from Rich's own interactive runs on the host he uses
most.
Host-level repairs made by hand (not by the script)
brew link rubyconflict, 5 hosts. The blocking file was<prefix>/lib/ruby/gems/4.0.0/gems/erb-6.0.7/libexec/erb, a hand-installed gem shim, byte-identical (sha256992f103c…) to the ruby 4.0.7 keg's own copy. Moved — not deleted — to~/.local/state/dev_update/archive/gem-link-conflicts-<ts>/with a README naming the restore command, thenbrew link ruby. Verifiedruby --versionanderb --versionstill resolve afterwards.keg tailscaleopted out, rdmpw3275m + rdmpw3265m./usr/local/bin/tailscaleis a 68-byte shim to/Applications/Tailscale.app/Contents/MacOS/Tailscale, and the app runs the daemon these hosts are actually on, so the shim must keep winning and the formula can never be linked. Recorded in~/.config/dev_update/self-heal-skip. Nothing depends on the formula; it was left installed rather than removed.- rdmpw3275m framework python.
/Library/Frameworks/Python.framework/Versions/3.14hadpydanticinstalled without its dependencies. Installedtyping-extensions,annotated-types,typing-inspection; upgradedpip,pydantic,pydantic_core.pip3 checknow returns "No broken requirements found";pbxprojstill imports. - rdmbair13m5. The 41h wedged dry run was killed (a
dry run mutates nothing). A separate orphaned
brew upgrade --greedy(ppid 1, 0% CPU, 1h42m) was cleared too and the interruptedbrew postinstall rubyrestarted.
Verification
Parsers unit-tested against real brew doctor /
brew upgrade --cask output with negative controls: an
unrelated cask error is not misclassified as a zombie receipt; opting
out a different keg does not hide a live one; a working
symlink, a plain file and a path outside the prefix are all left alone.
_fu_timed tested in both branches (124 on hang, output on
success, child exit code preserved). The wedged-lock gate tested in four
directions on rdmsm4x: old+ours cleared and lock taken · live+young
refused, holder alive · old+not-ours refused, unrelated pid alive · no
holder acquired cleanly.
Real runs, --no-system --no-mas, JSON summaries:
| Host | Issues before | Issues after |
|---|---|---|
| rdmsm4x | 1 (cask) + 1 warn (doctor) | 0 — brew doctor rc=0 "Your
system is ready to brew" |
| rdmpw3275m | 3 | 0 |
| rdmpw3265m | 1 | 0 |
| rdmbair15m5 | 1 | 0 — brew doctor "system is
ready to brew" |
| jdmbair13m5 | 2 | 1 (Xcode Apple ID sign-in — Rich only) |
| rdmbair13m5 | not running for 2 days | unblocked; catching up on 16 casks |
Files touched
~/dev/scripts/dev_update_v2.00.zsh(new),~/dev/scripts/dev_update.zsh(symlink repointed)- same file copied to
~/dev/scripts/on all five spokes, sha-verified, each spoke'sdev_update.zshsymlink repointed - new state, per host:
~/.local/state/dev_update/self-heal.tsv(ledger),~/.local/state/dev_update/archive/(archived conflicts),~/.config/dev_update/self-heal-skip(operator opt-outs; Intel hosts only so far)
How to undo
ln -sfn dev_update_v1.99.zsh ~/dev/scripts/dev_update.zshon any host — v1.99 is untouched.- Any archived file: the
README.txtbeside it names the exactmvthat restores it. brew unlink rubyreverses the linking;brew uninstall --cask <token>reverses a repair.
Open, filed, not fixed
ISSUE-20260920-04 — rdmbair13m5:
brew postinstall ruby never converges. Reproduced
twice and nowhere else on the fleet: the parent
brew.rb postinstall spins at ~100% CPU indefinitely while
the child postinstall.rb ruby sits idle at 0%. It holds
~188 Homebrew formula locks while it does, and it is why
ruby is still unlinked on that host and why two ruby kegs
(4.0.6_1 and 4.0.7) are present. The keg itself runs fine. Next step is
to sample the parent, not the child — the sample taken
today was of the idle child and showed only a normal
require stack.
brew uninstall --force ruby && brew install ruby is
the obvious next attempt but should be attended.
At the time of writing rdmbair13m5 also has an
interactive dev_update running in a Terminal window
(ttys013), stalled 15 minutes on a cask prompt nobody has answered. It
is bounded by STEP_TIMEOUT=3600 and will end itself; the
03:07 sweep arrives over ssh, takes the new unattended path, and is not
exposed to that prompt. No action needed.
Still needs Rich
- Developer ID
.p12is absent from the keychain on all five spokes. Exporting it needs the keychain password — Escalation Boundary #1. - jdmbair13m5: Xcode is not signed in to team ZU2882L4HT. Apple ID + 2FA.
- jdmbair13m5: macOS App Management for the terminal,
if
microsoft-edgeupgrades should stop erroring (they currently succeed anyway; the error is cosmetic).