Fleet changelogs · dev.ecs0.net
rdmbair15m5-changelog-20260906-0731-build39-recorded-booster-remediation-finished

rdmbair15m5 changelog — 2026-09-06 07:31 EDT — Build 39 recorded; priority-booster remediation finished; four claimed states re-verified

Session: claude@rdmbair15m5/dev-e5 · Window: 2026-09-06 07:16–07:31 EDT (this entry; the session itself began 2026-09-05 ~14:00 EDT) Budget posture: overall advice hold — OpenAI 7-day 100% (vendor-reported, authoritative), Anthropic 7-day 60% (conserve), 5-hour 3% (resets 13:50Z). All work below done inline; no fan-out, no delegation.


1. ReplicantDB Build 39 recorded in SESSION-STATE.md — 9cb964c

rdmsm4x:~/dev/apps/replicantDB/SESSION-STATE.md, 1880 → 1952 lines, ## headings 36 → 37 (asserted, not eyeballed). Pushed to fleet and backup; HEAD, fleet/main, backup/main all 9cb964c; git status --porcelain = 0.

Criteria were re-measured against dist-snapshots/1.19.5-build39/, not dist/ — dist/ is build scratch that the next build overwrites, so a criterion verified there proves nothing about the recorded artifact. All eight met: lipo x86_64 arm64 ×3 · minos 15.0 ×3 · version by execution 1.19.5 (39) "Restraint" · codesign --verify --deep --strict rc=0, Apple Development: Richard Doty (S65Q255HA8) → WWDR → Apple Root CA, TeamIdentifier=ZU2882L4HT · 16 ECS pins unmoved · 1206 tests / 0 failures (inherited — git diff 8c93766..3e641a8 -- Sources Tests = 0 files) · shasum -c 3 of 3 OK · hardened runtime confirmed absent (flags=0x0), as expected and deliberately not fixed.

The prior 05:23 entry, headlined ARTIFACT NOT CUT, was amended in place with a SUPERSEDED pointer rather than rewritten. Its failure analysis is what makes this build legible and stays intact; only the headline was out of date.

2. com.richh.priority.booster — the remediation was only half-applied. Now finished.

Previously recorded as "disabled". That was true but incomplete, and the gap mattered.

3. Four states claimed earlier, re-verified before being written down

Claim Verified state
agent_cred_hook.py write_to_env=True → False holds — lines 73 and 110 both write_to_env=False, header note present
fleet_load_canary_guard.py matcher tightened holds — line 136 exe == "ReplicantDB" and "--daemon" in args; two .bak files retained
1Password stopped, gui/501 agents disabled holds on rdmsm4x — all four agents disabled, process not running
~/.claude/CLAUDE.md oversight-vs-execution rule survived — the file is now v4.2 (2026-09-05), revised by another lane; the v4.1 rule is intact at line 53 and bypassPermissions is unchanged

4. Four probes of my own that measured the wrong thing — all caught before anything was recorded

Every one of these initially read as a finding. None was.

  1. grep -c '^#[A-Za-z_]*=' on global.env returned 0 commented keys. The actual form is # KEY=, with a space. Correct count: 7, exactly as recorded. The probe required no space and so could never have matched.
  2. pgrep -fl fleet_load_canary_guard → nothing, read as "the guard is not running". It runs on a 90-second StartInterval; the probe sampled the gap between cycles. It is healthy — last sweep 07:30:28, six hosts, exit 0.
  3. grep -i canary over LaunchAgents → nothing. The job's label is com.eastcoastscience.fleet-load-guard; "canary" appears only in the script's filename.
  4. ls -1 /Library/LaunchAgents/ /Library/LaunchDaemons/ | grep -i booster printed a filename that I attributed to the first directory. With two directory arguments ls -1 interleaves headers, and the file was in the second. A stat on the assumed path failed, which is what exposed it.

Same shape as the fleet's existing prior art (cmd | tail returning tail's status; ps -Ao command | grep -c matching its own shell): a probe that returns nothing because it asked the wrong question is indistinguishable from a clean result. In each case the correction came from checking the mechanism, not from re-running the probe.

5. rdmsm4x load, observed and resolved — not an incident

Load reached 41.49 at 07:31 on 16 cores, on the host that watchdog-panicked at 13:50 the previous day. Checked against the panic signature rather than assumed:

Guard behaviour was correct, and is worth knowing: is_high is ratio ≥ 2.5 **OR** load1 ≥ 40.0 for rdmsm4x. At 07:30:24 it measured ratio 2.24 / load 35.82 and logged OK — both conditions genuinely unmet. It would have fired at 41.49. The thresholds are doing what they were configured to do; no change made.

6. Also cleaned up

A find /Users/richh -iname '*priority*booster*' I had launched on rdmsm4x was still running in uninterruptible state against a ~2 M-item tree, and was the single D-state process on the host. Killed. Targeted directory checks answered the same question in seconds. My own contribution to the load I was investigating.


Not done in this pass, deliberately: ReplicantDB deployment, canary, daemon enablement, milestone 4 (S2/S3/S4/S8/S9), and the hardened-runtime fix. Build 38 remains deployed on 5 of 6 hosts; jdmbair13m5 is still blocked at 109 MB free (ISSUE-20260905-33 — Rich's personal data, untouched).

No secrets are recorded here. Key names and file locations only; values were never read into this record.