Fleet changelogs · dev.ecs0.net
rdmbair15m5-changelog-20260901-1311-syspolicyd-wedge-fixed-and-updater-v12

rdmbair15m5 changelog — 2026-09-01 13:00–13:11 EDT

Session 902def67 · run from rdmbair15m5 · continues the 12:56 changelog

What changed

Host Change
rdmbair13m5 sudo killall syspolicyd — daemon was wedged; launchd respawned it. claude-code runs again.
rdmbair13m5 com.apple.quarantine removed from the claude binary while testing a hypothesis. Did not help, not restored. Original value recorded in FINDINGS.
all six hosts ~/scripts/claude_code_update.zsh v1.1 → v1.2, mode 755, zsh -n clean. v1.1 backed up alongside.
rdmsm4x canonical ~/dev/fleet/claude-code-tcc-20260831/claude_code_update.zsh → v1.2; v1.1 in ~/.claude/backups/
rdmsm4x ~/dev/fleet/ISSUES.md — the rdmbair13m5 entry corrected and moved OPEN → RESOLVED; v1.2 entry added. 119 → 156 lines, OPEN count 21 → 20
rdmsm4x ~/dev/fleet/claude-code-tcc-20260831/FINDINGS-20260901.md — addendum appended, 188 → 276 lines

Root cause found: a wedged syspolicyd

claude --version on rdmbair13m5 sat at _dyld_start + 0, 0% CPU, no output, no exit — not one instruction of the program had run. Kernel log during a reproduction:

(AppleSystemPolicy) ASP: Sleep interrupted: ref 817, signal 0x100, pid: 94443
(AppleSystemPolicy) ASP: Security policy would not allow process: 94443, .../claude

Both lines appear at the moment of the SIGKILL, not before it. syspolicyd ran SecTrustEvaluateIfNecessary and nine SecTrustCopyAppleTrustAnchors calls, then went silent for twelve seconds. The process was never refused — it was waiting for a verdict that never arrived, and with no GUI session there was no dialog to render. Reading "would not allow" as a denial sends you after a Gatekeeper rule that does not exist.

syspolicyd was pid 528, up 19h03m — the only security daemon that had not restarted (amfid and trustd both restarted 4h13m earlier). Restarting it fixed it. spctl --status still assessments enabled developer id enabled — enforcement not weakened.

Heuristic worth keeping: a process parked at _dyld_start with 0% CPU and no output is a Gatekeeper verdict that never arrived. Go to syspolicyd first.

Three hypotheses falsified, so nobody re-runs them: quarantine xattr present (all six have it, five work); quarantine flag value (0181 on the broken host and two working ones); removing the xattr (still hung). Apple trust endpoints answered in 50–120 ms on all hosts.

Corrections to what I filed at 12:55 EDT

Updater v1.2

cc_version() — background + poll, deliberately not timeout(1) (coreutils, not guaranteed fleet-wide) — CC_TIMEOUT=45, exit 124 with a message naming syspolicyd, and a RECOVERED result when the pre-upgrade read times out but the post-upgrade read succeeds.

Verified three directions: hang → rc=124 empty; healthy → rc=0 + version; fast failure → rc=5, not conflated with the timeout. Executed live on rdmbair13m5 and rdmsm4x, both logging v1.2 UNCHANGED 2.1.252 -> 2.1.252.

Test trap worth more than the fix: the first run of that three-case test passed all three cases identically, because zsh script.zsh sources .zshenv, which reset PATH — all three "stubs" were silently the real claude. A test that cannot fail is not evidence. Use zsh -f and assert command -v resolves to the stub.

Still open, still needs Rich

No secrets recorded.