dev_update v1.99 — self-repair, permission-prompt mitigation, and a misdiagnosis fixed
- Host: rdmbair15m5 (Apple M5, macOS 27.0, arm64)
- Session start → end: 2026-09-12 03:34:43 EDT → 2026-09-12 04:45 EDT (~70 min)
- Agent: claude@rdmbair15m5 (Opus 5)
- Artifact:
~/dev/scripts/dev_update_v1.99.zsh, deployed to all six fleet Macs - sha256 (final build, identical on all six):
d2e12294097e682629e5948e865aaf3db3aecaeaaf5534bd5d617664401fcb3a
Why
Rich asked for the update run to act on its own findings
instead of only listing them: auto-parse the error log, take actions
(copy the developer identity, fix permissions, resolve errors),
optionally make an additional pass, and end with clear next steps or
detailed errors. Mid-task he added: mitigate the macOS permission
prompts that follow a brew upgrade of claude
and other applications.
The worked example was a Cascade Lake log ending:
██ HIGH ██ claude does not execute (rc=142): -- run: chmod +x /usr/local/Caskroom/claude-code@latest/2.1.269/claude
The three findings that shaped the design
That HIGH was a misdiagnosis.
rc=142is128+14= SIGALRM. Phase 3 wrapped the probe inperl -e 'alarm 15; exec @ARGV', so this was a timeout, not a permission fault. Worse,phase_claude_bitalready ranchmod +xon the binary before the probe — so the advice was both wrong and already done. Measured here: the binary is mode0755andclaude --versionreturns in 0.0 s.My first attribution was wrong, and I falsified it before shipping. I initially blamed
com.apple.quarantine: brew leaves it on the upgraded binary, so the first exec pays a Gatekeeper assessment of the whole 203 MB file. The controlled test on rdmpw3265m killed that theory — quarantine stripped, signature verified, correct x86_64 slice, andclaudestill hung past 180 s. Stripping quarantine is still worth doing (it removes a Gatekeeper re-check and prompts) but it is not the fix for this. See "The real finding" below, andISSUE-20260912-02.TCC grants split across two databases and only one is writable. Measured on this host: live grants pinned to versioned paths (
Cellar/zsh/5.9.2/bin/zsh,Cellar/herdr/0.8.2/bin/herdr) that the next upgrade will silently revoke, plus 4 dead rows (including a pre-rename/Users/rich/…path). These are in the SIP-sealed system db — not scriptable, not even as root. So the honest mitigation is prevention, priming and precise reporting, never a claim to have auto-granted.
What changed
Three new phases (the run is now 7 phases, not 5). The
embedded dev_update.zsh 2.13 payload was not
touched — verified byte-identical to v1.98 and to the shipped
constant 75d8e2d0298af30a….
| Phase | What it does |
|---|---|
| 5 permissions | Clears com.apple.quarantine from
signature-verified agent CLIs; audits TCC grants that
are dead or pinned to versioned paths; carries per-user-db grants
forward from a dead path to the live one; primes an FDA placeholder row
so a caller absent from System Settings becomes one click instead of
zero. |
| 6 remediation | Ingests dev_update's own
warnings[]/issues[] from the JSON (log only as
fallback), dispatches a rule table, and re-runs the original
probe after each fix. |
| 7 next steps | Always the last thing on screen, even on a clean run. CRIT/HIGH/WARN/INFO, each with why it was not automated and a paste-ready command. |
Fixers
Developer ID .p12 carried over from the basis host ·
mas · brew missing · the mechanical half of
brew doctor · xcode-select · Xcode first
launch · npm global PATH · ollama (127.0.0.1, Rule 23) · disk via
brew cleanup · tailscale. Findings with no fixer that name
their own remedy (… — run: mas upgrade) have that command
extracted and passed through rather than "go read the log".
Phase 3 now names the failure by its mechanism
142 timeout · 126 genuinely not executable
· 127 wrong architecture / missing dylib · other. Each
prints the measured discriminators — mode,
lipo -archs, codesign --verify, quarantine
state, size — and attributes nothing it has not tested. Default probe
timeout raised 15 s → 90 s (--claude-timeout).
Boundaries enforced in code
- Never automated: Apple ID / 2FA / keychain master password; spending money; destructive deletion of user data. These become NEXT STEPS with the exact command.
- TCC: only ever carries a grant forward from a dead path to the live path of the same binary — restoring a decision Rich already made. It never creates a grant that did not exist, and never claims to have written the sealed system db.
- Quarantine: CLI binaries only by default. ~30
/Applicationsbundles here carry the flag; they are counted and reported, never touched — clearing it machine-wide is a security-posture decision, not an update side effect.--fix-app-quarantineopts in. Nothing is cleared without a passingcodesign --verifyfirst. - Cleanup MOVES to
~/.local/state/fleet_update/archive/<what>-<ts>/. Nothing isrm'd. - The per-user TCC db is snapshotted to
~/.local/state/fleet_update/tcc-backups/before any write, and the snapshot is restored if the post-write re-read does not confirm.
The pass loop
Default is one pass. Re-running the whole updater is
the expensive way to re-check a fix (4m11s on a spoke, hours on Cascade
Lake); every fixer already re-runs its own original probe, which proves
the same thing for a fraction of the cost. --passes N
exists for the narrower case where a fix unblocks skipped work.
Two guards: a pass that fixes nothing ends the loop, and a pass whose
outstanding findings are identical to the previous pass's ends it and
says so.
Bugs found and fixed during the work
- A corrupted payload passed
--verify-embedded. An early edit landed inside the heredoc's comment header. The check still reported OK because the payload cache is keyed by the expected sha, so a corrupt heredoc still finds a good cache entry.--verify-embeddednow force-removes the cache so it validates this file. The corruption was reverted; payload proven byte-identical to v1.98. localinside a loop leaked variables to stdout. In zsh,local xfor a name already local in the same scope behaves like baretypesetand prints it. This putvout=''andcl='…'lines into the report. All declarations hoisted above their loops.- Multi-line next-step commands silently truncated the
JSON.
summary.kvis line-based, so a command spanning lines dropped every field after the break. Newlines are now encoded on the way out and decoded on the way in. ${${p:t:r}:l//…}— zsh rejects chaining:lwith//in one expansion ("unrecognized modifier"). Split into three steps. The same block also calledbrew list --caskonce per app (~90 invocations); the token list is now resolved once.- The
.p12transport passphrase.security export -Pputs the passphrase in argv, visible in the basis host's process table. The long-livedFLEET_SIGNING_P12_PASSWORDis no longer what sits there: the in-flight.p12is wrapped with a one-time random passphrase, worthless after import. The fleet secret is used only locally. - The remote export could hang forever on a
SecurityAgent dialog on the basis host's console
(
ConnectTimeoutbounds only the handshake). Now capped at 45 s — exceeding it is a result, reported as such, per the patience budget. - Exit code and summary banner contradicted each other ("0 items need attention", exit 20). Both now report the post-remediation picture; a finding that was fixed and re-verified no longer colours the exit code.
Verification (counts, not "green")
zsh -nclean;--verify-embeddedOK against a cold cache on both hosts.- Payload sha256
75d8e2d0298af30a…— byte-identical across v1.98, v1.99 and the constant. - Full 7-phase dry run on rdmbair15m5: all phases executed, exit 20, NEXT STEPS rendered.
- Live proof of the quarantine fix, phase 3:
quarantine cleared (signature verified)→claude --version … 2.1.269 (0s). - Fed the exact findings from Rich's Cascade Lake log: the signing fixer attempted the export, was refused by the keychain ACL as predicted, and degraded to an exact-command NEXT STEP. Xcode account correctly escalated, never attempted.
- JSON gains
remediation_seen,remediation_fixed,remediation_fixed_items,remediation_unfixed,next_steps_n, and structurednext_steps[](severity/title/why/command/where). Confirmed 2 rows with 2- and 8-line commands.
The
real finding: claude 2.1.269 does not start on 4 of 6
Macs
Filed as ISSUE-20260912-02 (BUG, high).
Measured 2026-09-12 ~04:10 EDT:
| Host | claude --version |
|---|---|
| rdmbair15m5, rdmsm4x | rc=0 in 0 s |
| rdmbair13m5, jdmbair13m5, rdmpw3265m, rdmpw3275m | never returns — rc=142 at
20/25/120/180 s |
Ruled out by direct measurement: exec bit (-rwxr-xr-x),
architecture (correct slice), code signature (verifies, Q6L2SF6YDW),
quarantine (stripped, still hangs), user config (clean HOME
still hangs), concurrent claude process count (4 hangs, 8
works), the binary itself (byte-identical 203,150,240 across arm64 —
works on two, hangs on two), and SSH vs local context (rdmbair15m5
returns in 0 s both ways).
sample shows only _dyld_start at 0 % CPU
with amfid and syspolicyd idle — but a 200 MB single-file bundle does
not symbolicate, so that is unresolved frames, not evidence of a
dyld block. I over-read it at first and have corrected that
here.
What v1.99 does about it: phase 3 no longer guesses.
It prints the measured discriminators
(mode … arch … signature … quarantine … bytes), says the
process blocks before main() so no local permission change
fixes it, and files a NEXT STEP with the commands to re-confirm and to
fall back to the previous cask version. Two wrong diagnoses have now
been made on this symptom — v1.98's "run: chmod +x" and my own
quarantine theory — so the script states what it measured and attributes
nothing it has not tested.
Deployment state
rdmbair15m5:~/dev/scripts/dev_update_v1.99.zsh— symlinksdev_update.zsh,devupdate.zsh,fleet_update.zsh,fleet_upgrade.zsh,updatedev.zshrepointed.rdmsm4x:~/dev/scripts/dev_update_v1.99.zsh— same, verified there.- v1.98 left in place on both hosts. Rollback is one
command:
for l in dev_update devupdate fleet_update fleet_upgrade updatedev; do ln -sfn dev_update_v1.98.zsh ~/dev/scripts/$l.zsh; done - Deployed to all six hosts, each verified by sha256
match,
zsh -n,--verify-embeddedagainst a cold cache, and-V. Reference sha of the final build:d2e12294097e6826…(the file was revised after the first push; every host carries the final one).
Further bugs found after the first deploy
-Vhad drifted — it printed "dev_update_v1.96.zsh … dev_update 2.11" while carrying 2.13. Now derived fromFL_VERSION/DU96_DU_VERSION/ the payload sha, and emitted after those constants exist (inline in the arg loop yielded empty fields).- A
|in a next-step command truncated the field.NS_ITEMSpacked five fields with|, andsample $! 3 -mayDie | head -20; kill $!lost everything after-mayDie. Fields now use US (0x1f). - A
;in a command truncated the record — same defect one level up, since records were joined with;and the command containsls -d … ; # then run …. Records now use RS (0x1e). - zsh
(j:\x1e:)joins with the literal four characters, not the byte. Needs thepflag:(pj:\x1e:). Caught because the JSONwherefield readrdmpw3265m\x1eINFO. Round-trip now verified on Intel: 3 declared, 3 parsed, 6-line command with its pipe intact.
Revision pass (2026-09-12 04:45 EDT) — stale self-references and a real leak
The file had accumulated version drift across five prior releases.
All of it now reads 1.99, except the changelog entries, which
are history and were deliberately left alone. Every edit was
asserted to fall outside the payload heredoc, and the payload was
re-proven byte-identical (75d8e2d0298af30a…)
afterwards.
- Header title,
Version:line,Run:/Log:lines:v1.98→v1.99, and theLog:path corrected — it namedfleet_update-<ts>.log, which the code has not written for releases. Usage:, the root-refusal message, the unknown-option message, the payload manifest and its three error paths:dev_update_v1.96.zsh→dev_update_v1.99.zsh.- The sudo prompt still said
dev_update_v1.93. - The status strip registered as
dev_update_v195and titled itselfdev_update v1.95; the title is now$FU_VERSION, so it cannot drift again. - Two lines stated wrong facts, not just stale names:
--helpadvertised "footlights.zsh 1.0.0 + dev_update.zsh 2.10" and the payload banner said 2.11, while the carried payload is 2.13 and footlights is 1.1.0. - The private-copy temp name
(
.dev_update_v195.pinned.$$.zsh) is functional — one site creates it and two traps delete it by pattern. All three renamed together; a mismatch would have leaked a 275KB copy per run silently.
The leak that the rename exposed
Checking the three sites found 12 orphaned pinned copies,
3.2MB, from this session alone — and two of them were created
by the -V and --verify-embedded commands I had
just run. Cause: the copy is removed by the EXIT trap, but the trap is
installed late, so every early-exit path
(-V, --verify-embedded, --help,
an unknown option) leaked one, as does any run killed before the trap
arms.
Two fixes: _fu_unpin on each early-exit path, and
_fu_sweep_pinned at startup, which removes orphans
keyed on the PID embedded in the filename. A copy whose
PID is still alive is never touched — that process is reading it by byte
offset and deleting it would crash the run. Verified in both directions:
all four early-exit paths and a full run now leak zero, all 12
pre-existing orphans were swept, and a planted copy belonging to a live
PID survived the sweep.
Re-verified on all six hosts after the revision pass: sha match,
zsh -n, payload verify against a cold cache, symlinks 5/5,
0 orphans, and a full 7-phase run on rdmpw3265m (exit 20,
next_steps declared 3 / parsed 3).
Outstanding on rdmbair15m5 (from the run's own NEXT STEPS)
zshholds FDA + Accessibility onCellar/zsh/5.9.2/bin/zsh; the next upgrade revokes both. Re-grant/opt/homebrew/bin/zshinstead.- 4 stale TCC rows in the sealed system db (clutter, not a fault).
- 8 cask app bundles carry
com.apple.quarantine— untouched by design. - The Developer ID identity is still absent here; the export needs one keyboard-present step on rdmsm4x (keychain ACL).
No secrets are recorded in this file.
FLEET_SIGNING_P12_PASSWORD is referenced by name
and location (~/.secrets/global.env) only.