rdmsm4x-changelog-20260831-1934-claude-code-2.1.252-fleet-rollout-and-syspolicyd-wedge
Window: 2026-08-31 18:46:32 – 19:34:51 EDT (~48 min)
· Lead: claude@rdmsm4x session dev-94 [785f8b]
One-line summary: brought all six Macs to Claude Code 2.1.252,
deployed the missing claude_code_update auto-update job to
the four hosts that never had it, and cleared a wedged
syspolicyd on jdmbair13m5 that was deadlocking every exec
of the claude binary.
Scope
All six fleet hosts: rdmsm4x, rdmbair15m5,
rdmbair13m5, jdmbair13m5,
rdmpw3265m, rdmpw3275m.
What changed
1. Claude Code upgraded on rdmsm4x
brew upgrade --cask claude-code@latest— 2.1.251 → 2.1.252 (18:46:59).- Asserted the exec bit per the documented post-upgrade mode-644 trap; it was already 755.
2. Auto-update mechanism deployed to 4 hosts (the real gap)
The version numbers looked fine fleet-wide, but the
mechanism keeping them current existed on only 2 of 6 hosts.
rdmbair15m5, rdmbair13m5,
jdmbair13m5 and rdmpw3275m had
neither ~/scripts/claude_code_update.zsh
nor the launchd job — they were current by coincidence,
with nothing to keep them current.
Deployed to those four, byte-identical to the rdmsm4x copy:
~/scripts/claude_code_update.zsh— v1.1, sha256 prefix9a586985faf9, mode 755~/Library/LaunchAgents/com.eastcoastscience.claude-code-update.plist(StartInterval21600 = 6h,RunAtLoad), generated per-host so log paths use the real$HOMElaunchctl bootstrap gui/$(id -u)+kickstart— usedid -u, not a hardcoded 501, becauserichhis UID 502 on jdmbair13m5 (documented non-uniform UID;juliaholds 501 there).
Left alone: rdmsm4x and rdmpw3265m already
had v1.1 at the same sha.
3. jdmbair13m5 — wedged syspolicyd (the actual incident)
Symptom: claude --version hung
indefinitely (>7 min, rc=124 at every timeout). Binary was mode 755,
byte-identical (sha256 b661c6a094fcc32656bf7c0071c5b45b,
197,220,928 bytes) to two hosts where it ran fine.
Ruled out by measurement, not assumption:
- Not the binary — sha256 identical across jdmbair13m5 / rdmbair13m5 / rdmsm4x.
- Not the exec bit — mode 755.
- Not arch — arm64 on arm64; signature and Developer ID identical to a working host.
- Not quarantine — a working host carried the identical
com.apple.quarantinevalue. - Not config — an isolated
CLAUDE_CONFIG_DIRhung the same way. - Not resources — load 2.63, 91% memory free, zero swap, 143 GB free; the 197 MB binary read end-to-end in 0.005 s.
Root cause, from spindump:
AppleSystemPolicy::procNotifyExecComplete
AppleSystemPolicy::waitForEvaluation
ASPEvaluationManager::waitOnEvaluation
lck_mtx_sleep (indefinite)
The AppleSystemPolicy kernel extension blocks exec until
syspolicyd returns a Gatekeeper verdict.
syspolicyd (pid 532) had stopped servicing its evaluation
queue, so the kernel slept forever. The tell that misleads:
syspolicyd showed 0.0% CPU and the syspolicy/AMFI log was
silent — it looked healthy precisely because it was wedged. An
idle security daemon is not evidence of an idle security path.
Fix: sudo killall syspolicyd (launchd
respawned it as pid 45627). claude --version returned
2.1.252 rc=0 immediately after. No reboot needed.
Verification (end state, all six)
Every host re-measured after the work — version by actually running the binary, exit code captured separately rather than through a pipe:
| Host | claude | rc | mode | script | job | last run log |
|---|---|---|---|---|---|---|
| rdmsm4x | 2.1.252 | 0 | 755 | v1.1 9a586985faf9 |
loaded exit=0 | UNCHANGED 2.1.252 |
| rdmbair15m5 | 2.1.252 | 0 | 755 | v1.1 9a586985faf9 |
loaded exit=0 | UNCHANGED 2.1.252 |
| rdmbair13m5 | 2.1.252 | 0 | 755 | v1.1 9a586985faf9 |
loaded exit=0 | UNCHANGED 2.1.252 |
| jdmbair13m5 | 2.1.252 | 0 | 755 | v1.1 9a586985faf9 |
loaded exit=0 | UNCHANGED 2.1.252 |
| rdmpw3265m | 2.1.252 | 0 | 755 | v1.1 9a586985faf9 |
loaded exit=0 | UNCHANGED 2.1.252 |
| rdmpw3275m | 2.1.252 | 0 | 755 | v1.1 9a586985faf9 |
loaded exit=0 | UNCHANGED 2.1.252 |
Outstanding / owner actions
- jdmbair13m5 quarantine xattr — REMOVED THEN RESTORED
(2026-08-31 19:45 EDT). I removed
com.apple.quarantinewhile testing; it was not the cause. Restored at Rich's direction for fleet consistency, and verified byte-identical by hex to an untouched host's form:0181;6a9604dd;Homebrew\x20Cask;2894C019-0401-4BCA-A685-537C7E09952Awhere\x20is the four literal characters backslash-x-2-0, NOT a space. My first restore attempt wrote a real space (0x20) and was wrong; corrected viaxattr -wxwith explicit hex.claude --versionreturns 2.1.252 rc=0 with the flag in place. - Note on what quarantine means here: it is Homebrew
Cask's default provenance stamp on every cask it installs, not a malware
finding. 50 of 96 apps in /Applications on rdmsm4x carry it (1Password,
Bitwarden, ChatGPT, Microsoft Edge, fonts). The binary itself assessed
as
source=Notarized Developer ID,accepted, Developer ID Application: Anthropic PBC (Q6L2SF6YDW),codesign --verify --strictclean. - Fleet inconsistency noticed, not acted on:
rdmbair15m5's quarantine value has an EMPTY agent field
(
0181;6a9604c5;;EC3134A0-...) where other hosts carryHomebrew\x20Cask. Harmless; recorded only so a future audit does not read it as tampering. - Why syspolicyd wedged is not established. Host had been up 1h26m. If it recurs, that is a pattern worth a real investigation rather than another killall.
- The stale
read-only variable: statusline inclaude_code_update.launchd.logon rdmsm4x and rdmpw3265m is historical (the v1.0 zsh bug that v1.1 fixed), not a live fault.
How to undo
- Auto-update job:
launchctl bootout gui/$(id -u)/com.eastcoastscience.claude-code-updateand remove the plist +~/scripts/claude_code_update.zshon the four hosts. - Claude Code version:
brew uninstall --cask claude-code@latestthen reinstall a pinned version. - syspolicyd restart is self-healing and needs no undo.
No secrets were read or written.