Fleet changelogs · dev.ecs0.net
rdmbair15m5-changelog-20260912-0352-dev-update-v199-self-repair-and-tcc-prompt-mitigation

dev_update v1.99 — self-repair, permission-prompt mitigation, and a misdiagnosis fixed


Why

Rich asked for the update run to act on its own findings instead of only listing them: auto-parse the error log, take actions (copy the developer identity, fix permissions, resolve errors), optionally make an additional pass, and end with clear next steps or detailed errors. Mid-task he added: mitigate the macOS permission prompts that follow a brew upgrade of claude and other applications.

The worked example was a Cascade Lake log ending:

 ██ HIGH ██  claude does not execute (rc=142):  -- run: chmod +x /usr/local/Caskroom/claude-code@latest/2.1.269/claude

The three findings that shaped the design

  1. That HIGH was a misdiagnosis. rc=142 is 128+14 = SIGALRM. Phase 3 wrapped the probe in perl -e 'alarm 15; exec @ARGV', so this was a timeout, not a permission fault. Worse, phase_claude_bit already ran chmod +x on the binary before the probe — so the advice was both wrong and already done. Measured here: the binary is mode 0755 and claude --version returns in 0.0 s.

  2. My first attribution was wrong, and I falsified it before shipping. I initially blamed com.apple.quarantine: brew leaves it on the upgraded binary, so the first exec pays a Gatekeeper assessment of the whole 203 MB file. The controlled test on rdmpw3265m killed that theory — quarantine stripped, signature verified, correct x86_64 slice, and claude still hung past 180 s. Stripping quarantine is still worth doing (it removes a Gatekeeper re-check and prompts) but it is not the fix for this. See "The real finding" below, and ISSUE-20260912-02.

  3. TCC grants split across two databases and only one is writable. Measured on this host: live grants pinned to versioned paths (Cellar/zsh/5.9.2/bin/zsh, Cellar/herdr/0.8.2/bin/herdr) that the next upgrade will silently revoke, plus 4 dead rows (including a pre-rename /Users/rich/… path). These are in the SIP-sealed system db — not scriptable, not even as root. So the honest mitigation is prevention, priming and precise reporting, never a claim to have auto-granted.

What changed

Three new phases (the run is now 7 phases, not 5). The embedded dev_update.zsh 2.13 payload was not touched — verified byte-identical to v1.98 and to the shipped constant 75d8e2d0298af30a….

Phase What it does
5 permissions Clears com.apple.quarantine from signature-verified agent CLIs; audits TCC grants that are dead or pinned to versioned paths; carries per-user-db grants forward from a dead path to the live one; primes an FDA placeholder row so a caller absent from System Settings becomes one click instead of zero.
6 remediation Ingests dev_update's own warnings[]/issues[] from the JSON (log only as fallback), dispatches a rule table, and re-runs the original probe after each fix.
7 next steps Always the last thing on screen, even on a clean run. CRIT/HIGH/WARN/INFO, each with why it was not automated and a paste-ready command.

Fixers

Developer ID .p12 carried over from the basis host · mas · brew missing · the mechanical half of brew doctor · xcode-select · Xcode first launch · npm global PATH · ollama (127.0.0.1, Rule 23) · disk via brew cleanup · tailscale. Findings with no fixer that name their own remedy (… — run: mas upgrade) have that command extracted and passed through rather than "go read the log".

Phase 3 now names the failure by its mechanism

142 timeout · 126 genuinely not executable · 127 wrong architecture / missing dylib · other. Each prints the measured discriminators — mode, lipo -archs, codesign --verify, quarantine state, size — and attributes nothing it has not tested. Default probe timeout raised 15 s → 90 s (--claude-timeout).

Boundaries enforced in code

The pass loop

Default is one pass. Re-running the whole updater is the expensive way to re-check a fix (4m11s on a spoke, hours on Cascade Lake); every fixer already re-runs its own original probe, which proves the same thing for a fraction of the cost. --passes N exists for the narrower case where a fix unblocks skipped work. Two guards: a pass that fixes nothing ends the loop, and a pass whose outstanding findings are identical to the previous pass's ends it and says so.

Bugs found and fixed during the work

Verification (counts, not "green")

The real finding: claude 2.1.269 does not start on 4 of 6 Macs

Filed as ISSUE-20260912-02 (BUG, high). Measured 2026-09-12 ~04:10 EDT:

Host claude --version
rdmbair15m5, rdmsm4x rc=0 in 0 s
rdmbair13m5, jdmbair13m5, rdmpw3265m, rdmpw3275m never returns — rc=142 at 20/25/120/180 s

Ruled out by direct measurement: exec bit (-rwxr-xr-x), architecture (correct slice), code signature (verifies, Q6L2SF6YDW), quarantine (stripped, still hangs), user config (clean HOME still hangs), concurrent claude process count (4 hangs, 8 works), the binary itself (byte-identical 203,150,240 across arm64 — works on two, hangs on two), and SSH vs local context (rdmbair15m5 returns in 0 s both ways).

sample shows only _dyld_start at 0 % CPU with amfid and syspolicyd idle — but a 200 MB single-file bundle does not symbolicate, so that is unresolved frames, not evidence of a dyld block. I over-read it at first and have corrected that here.

What v1.99 does about it: phase 3 no longer guesses. It prints the measured discriminators (mode … arch … signature … quarantine … bytes), says the process blocks before main() so no local permission change fixes it, and files a NEXT STEP with the commands to re-confirm and to fall back to the previous cask version. Two wrong diagnoses have now been made on this symptom — v1.98's "run: chmod +x" and my own quarantine theory — so the script states what it measured and attributes nothing it has not tested.

Deployment state

Further bugs found after the first deploy

Revision pass (2026-09-12 04:45 EDT) — stale self-references and a real leak

The file had accumulated version drift across five prior releases. All of it now reads 1.99, except the changelog entries, which are history and were deliberately left alone. Every edit was asserted to fall outside the payload heredoc, and the payload was re-proven byte-identical (75d8e2d0298af30a…) afterwards.

The leak that the rename exposed

Checking the three sites found 12 orphaned pinned copies, 3.2MB, from this session alone — and two of them were created by the -V and --verify-embedded commands I had just run. Cause: the copy is removed by the EXIT trap, but the trap is installed late, so every early-exit path (-V, --verify-embedded, --help, an unknown option) leaked one, as does any run killed before the trap arms.

Two fixes: _fu_unpin on each early-exit path, and _fu_sweep_pinned at startup, which removes orphans keyed on the PID embedded in the filename. A copy whose PID is still alive is never touched — that process is reading it by byte offset and deleting it would crash the run. Verified in both directions: all four early-exit paths and a full run now leak zero, all 12 pre-existing orphans were swept, and a planted copy belonging to a live PID survived the sweep.

Re-verified on all six hosts after the revision pass: sha match, zsh -n, payload verify against a cold cache, symlinks 5/5, 0 orphans, and a full 7-phase run on rdmpw3265m (exit 20, next_steps declared 3 / parsed 3).

Outstanding on rdmbair15m5 (from the run's own NEXT STEPS)

  1. zsh holds FDA + Accessibility on Cellar/zsh/5.9.2/bin/zsh; the next upgrade revokes both. Re-grant /opt/homebrew/bin/zsh instead.
  2. 4 stale TCC rows in the sealed system db (clutter, not a fault).
  3. 8 cask app bundles carry com.apple.quarantine — untouched by design.
  4. The Developer ID identity is still absent here; the export needs one keyboard-present step on rdmsm4x (keychain ACL).

No secrets are recorded in this file. FLEET_SIGNING_P12_PASSWORD is referenced by name and location (~/.secrets/global.env) only.