jdmbair13m5 changelog — xreaper zsh special-variable bugs + fleet alias unshadow
Session: 2026-09-04 04:38:35 → 05:13:20 EDT (~35
min) · run from jdmbair13m5 Asked for:
remove fd from all .zshrc files fleet-wide so
normal shell commands run; fix ~/dev/apps/xreaper/scripts/
(canonical source on rdmsm4x) and every fleet copy.
The reported symptom
▉ INFO gone ......... sshd-session: richh [priv] (2486) exited before we reached it
base='sshd-session: richh [priv]'
SIGSTOP codex (3487) -- suspended, cpu relieved, nothing lost yet
_xr_capture_codex:local:8: path: inconsistent type for assignment
Two independent defects in one paste, and diagnosing them uncovered two more.
Root causes — four bugs, all reproduced before being fixed
BUG 1 — hard error. xreaper.zsh:1102,
_xr_capture_codex:
local mt="${c%% *}" path="${c#* }"path is a zsh special parameter, tied
as an array to $PATH. The function aborts at that line —
local:8 is literally the 8th line of the function.
Codex and Claude session capture never worked. Fixed by
renaming the loop variable path → xrp (and its
two ${path:t:r} readers).
BUG 2 — why only codex failed. One line above:
local -a live c declares the loop scalar
c as an array, so ${c#* } is an array-valued
expansion — which is exactly what makes
local path=<that> a type conflict.
_xr_capture_claude declares local c
separately, so it survived. Verified both ways in isolation: the same
statement is legal with c scalar and fatal with
c array.
BUG 3 — the base=... noise. Not a
leftover debug statement. zsh's TYPESET_SILENT is
off by default, and local NAME with no
value prints the parameter when it already exists
in scope. local base; base="$(xr_base "$comm")" sits inside
two per-process loops, so from the second iteration on, all ~950
processes each echoed a line into the report. The quoting in
base='sshd-session: richh [priv]' is zsh's own typeset
format — that is the tell. Fixed with setopt TYPESET_SILENT
beside the existing setopt NO_NOMATCH.
BUG 4 — silent wrong answer.
XR_CAP_CONF=$(( ${#live} == 1 )) && XR_CAP_CONF="exact" || XR_CAP_CONF="probable"The assignment returns 0 regardless of the arithmetic, so the
&& branch always won. Capture confidence was
reported exact unconditionally. Fixed with a real
if/else.
.zshrc — the alias
shadowing
Removed only aliases that shadow a real command with a
different tool: find=fd, grep=rg,
cat=bat, ls=eza. Nothing lost —
fd, rg, bat, eza all
remain callable under their own names.
ll/la/lt kept (new names, not
overrides).
Deliberately KEPT: oh-my-zsh's
alias grep='grep --color=auto ...' and
alias ls='ls -G'. Those are the same program, tweaked —
removing them would be scope creep, not a fix. The operative rule:
an alias is a shadow only when the first word of its value differs
from its name.
This also closes the trap documented in ~/dev/CLAUDE.md
§9: bare ls → bare eza prints nothing
and exits 0 in a non-interactive shell — an empty listing that
reads as a true answer.
Tooling (idempotent, kept for reuse)
~/dev/scripts/xreaper-fix-20260904/
xr_fix.pyv1.4 +xr_fix.zsh— the four xreaper bugszshrc_unshadow.pyv2.0 +zshrc_unshadow.zsh— the alias shadowing
Both back up to <file>.bak-<timestamp>,
assert every intended change landed and that nothing
offending survived, and gate on zsh -n
before installing. An already-correct file reports
CLEAN.
Two defects found in that tooling by the fleet run, and fixed
- v1.3 aborted on build fragments
(
xreaper.body/tail.zsh) — they carry nosetopt NO_NOMATCHline to anchor bug 3 to, and the assertion fired, reporting "transform produced nothing". Anchor absence is now a no-op. - v1.0 of the zshrc patcher matched only single-quoted,
one-per-line aliases. Some fleet hosts chain them:
... && alias ls="eza" && alias ll="..." && alias la="...". Commenting the whole line would have silently takenll/la/ltwith it. v2.0 removes only the offending link and rebuilds the line. It also stopped a false FAIL caused by two disagreeing definitions of "shadow".
Convergence
Every patched host produces a byte-identical file:
| file | unpatched | patched |
|---|---|---|
xreaper.zsh |
17a491eed2f697c6c2190aaf3fd7c483 |
156e2d1175a1f99c6ee4effb47866717 |
xreaper.body.zsh |
d165e979f3eee82e8e8d9b74c37faf27 |
b0192d3a17e944ba54abd751868eeaa0 |
xreaper.tail.zsh |
39a4e1867a61949d5993a9c089b110a3 |
fd4fecc2861f8e7db0fbb4d157692095 |
reap.zsh |
9de073550ae31f5f874f09876cf0def0 |
unchanged (unrelated, correctly untouched) |
Per-host results
| host | xreaper files | .zshrc shadows removed |
dry-run verification |
|---|---|---|---|
jdmbair13m5 (arm64, 27.0) |
3 fixed | 4 | 946 procs · 9 findings · 0 errors · 0 leaks |
rdmpw3275m (Intel, 26.7) |
3 fixed | 2 | 273 procs · 8 findings · 0 errors · 0 leaks |
rdmbair13m5 (arm64) |
6 fixed, 2 locations | 5 | 914 procs · 11 findings · 0 errors · 0 leaks |
rdmbair15m5 (arm64) |
3 fixed | (already clean) | 865 procs · 7 findings · 0 errors · 0 leaks |
rdmsm4x (canonical) |
6 fixed, 2 locations | 2 | 343 procs · 9 findings · 0 errors · leak 699 → 0 |
rdmpw3265m (Intel) |
3 fixed | 2 | 268 procs · 8 findings · 0 errors · 0 leaks |
The canonical directory is apps/xReaper —
capital R, not apps/xreaper.
rdmbair13m5 carries the same second copy. On every host
with two copies they are byte-identical.
rdmsm4x specifics (the canonical host)
- Committed:
2f4b066fix(xreaper): repair four zsh defects in the agent-session capture path, by explicit path — verified to touch exactlyscripts/xreaper.{zsh,body.zsh,tail.zsh}, 3 files, +20/-14. Nogit add -A. - Test suite: 22 passed, 0 failed. Record the number, not the word "pass" — a later drop from 22 is what would reveal a silent revert; "tests pass" cannot detect a deletion.
- BUG 1 proved directly, not by absence. A dry run
never calls the capture path, so "0 errors in a dry run" is NOT evidence
the fix worked.
_xr_capture_codexwas extracted and run standalone against a live codex pid: it now completes and returns a real resumable session id (codex resume 01a06aab-…). Before the fix it aborted at line 8. - BUG 3 measured both ways on a real run: 699
leaked
base=lines → 0, report length 771 → 69 lines. --verify-embedded→ ALL CLEAR, universal2x86_64 arm64arch gate passes.rc=20is not a failure:_xr_exit_code()returns 20 when the reaped list is non-empty, 30 on real failures. Two zombie-parent findings in dry-run ⇒ 20.
Durability gap found on rdmsm4x — worth acting on separately
~/dev/apps/xReaper has no git remote.
git remote -v is empty (verified directly), so the repo is
local-only: not mirrored to git.ecs0.net and absent from
FLEET-PROJECT-POINTERS.md. On the canonical host that means
the only copy of this project's history is on one machine. Out of scope
for this fix; flagged, not silently left.
Notes for the next session
- No host touched by this work is a git repository
(
git rev-parsefails at every level), so nothing was committed anywhere. The canonical copies still need promoting throughrdmsm4x/git.ecs0.net. rdmpw3275mfirst probed as dead (no ICMP, port 22 filtered, no tailscale reply) but was reachable minutes later. Fleet SSH needsConnectTimeout=45— short timeouts fail during the handshake and read as "host down". Three hosts were misdiagnosed that way in this session before the longer timeout was tried.- Budget posture at fan-out was
conserve(Anthropic 7-day at 71%, vendor-reported).
Final convergence — verified directly, host by host, after all agents reported
20 files across 6 hosts, every basename byte-identical everywhere, 0 live shadowing aliases on all 6:
jdmbair13m5 dev/scripts/{xreaper.zsh, .body.zsh, .tail.zsh} shadows=0
rdmsm4x dev/apps/xReaper/scripts/* + dev/scripts/* (6 files) shadows=0
rdmbair13m5 dev/apps/xReaper/scripts/* + dev/scripts/* (6 files) shadows=0
rdmbair15m5 dev/scripts/* (3 files) shadows=0
rdmpw3265m dev/scripts/* (3 files) [Intel] shadows=0
rdmpw3275m dev/scripts/* (3 files) [Intel] shadows=0
Both Intel Macs built the embedded rusage_probe.c
cleanly and reported PASS rusage_probe: cached. No
Intel-specific lipo -archs problem exists on that pair.
Method note worth keeping
A dry run never calls the capture path, so "0
inconsistent type errors in a dry run" is NOT evidence that
BUG 1 was fixed — the faulty line simply never executed. The proof came
from extracting _xr_capture_codex and running it standalone
against a live codex pid on rdmsm4x: it completes and
returns a real resumable session id. This is the same failure shape as
everything in ~/.claude/CLAUDE.md § shared trees — an
absence that reads as a pass. Bare eza returning
nothing with rc=0, a green build after a silent revert, and a clean dry
run over a never-executed branch are all the same defect class.