rdmsm4x - fix - superseded load guard refuses + HostPressureLevel doc contract - claude - ISSUE-20260924-35 - tyrell-fleet - 20260924-2215
Session: claude@rdmsm4x/879ddeaa · 2026-09-24 22:05:27 - 22:15:08 EDT Trigger: "resume all work" (relayed by peer dev-c2, with budget=hold and a no-launchd/no-TCC request) Summary: Verified my own 2026-09-05 load-grading finding was genuinely fixed in Tyrell (it was, and is working on live data), then closed the two gaps the fix left behind: a public doc comment still stating the pre-fix contract, and a superseded script that could still kill daemons on the bad metric.
Scope
Host: rdmsm4x only. No remote host touched. No launchd, no TCC/FDA, no xcode-select — peer dev-c2 asked the fleet to leave those alone tonight while priming Full Disk Access for ECSToolHost.
What changed
1.
~/scripts/fleet_load_canary_guard.py — now refuses to run
(local state change)
Added a module-level refusal after the imports: exits
2 with a message naming the successor (Tyrell
FleetGuardCoordinator), why the old grading is wrong, and
the override FLEET_LOAD_GUARD_ALLOW_SUPERSEDED=1.
Exit 2 deliberately, matching the fleet contract — 0 pass / 1 fail / 2 could-not-be-evaluated. A superseded tool is the third thing, not the first two.
Why it mattered more than a stale file: the script's
own docstring says it stops the ReplicantDB canary daemon, kills test
runners (swift-test, xctest,
swiftpm-testing-helper) and broadcasts slowdown directives
to every agent — all graded on load1/ncpu alone. Live
reading on this host at 2026-09-25T02:01:42Z: load1
105.07, ratio 6.57, cpuIdlePct 57. It would have called that a
storm and started killing, while the CPU was 57% idle.
Verified both directions:
| case | result |
|---|---|
| default run | rc=2, refusal printed, no action taken |
FLEET_LOAD_GUARD_ALLOW_SUPERSEDED=1 |
rc=0, executes past the guard |
python3 -m py_compile |
OK |
The override case is the one that matters — it proves the script refuses rather than being broken.
2.
Tyrell HostPressureLevel doc comments — PR #45 (branch, not
canonical)
ISSUE-20260924-11 added the CPU-idle gate and documented
it at the function, but the public enum's per-case comments still read
High (load ratio >= 2.0) /
Critical (load ratio >= 3.0) with no mention of it.
Those render in autocomplete and generated docs, so the API surface
still advertised the contract the fix deliberately stopped honouring.
Comments only; evaluate() and every threshold
untouched.
Each claim checked against the code, not paraphrased:
thermal-critical returns before the idle branch (so it does
override idle entirely), and the idle branch's highest return is
.elevated (so Elevated really is the ceiling).
swift build --target TyrellCore clean;
HostPressureEvaluatorTests 6/6.
Verification that the original finding actually shipped
Not taken from ticket status — measured.
~/.agent-coordination/fleet-load-guard-status.json is now
written by Tyrell (FleetGuardCoordinator.swift:455) and
carries cpuBusyPct, cpuIdlePct,
uninterruptibleWaitCount. The sample above IS the condition
the original ticket described, and
HostPressureLevel.evaluate now caps it at
.elevated instead of .critical.
A correction I made to my own filing
I first recorded "could not verify there is no caller" after two
greps were killed at 120s/60s. They completed in the background
afterwards: the only live reference is the dormant
plist; every other hit is a read-only 2026-09-05 snapshot or archived
ticket HTML. Positive control held (44 of 58 plists match
eastcoastscience). A zero from a killed grep and a zero
from a completed one look identical, so the weaker claim should not have
been left standing.
How to undo
- Script: delete the guarded block after the imports (it is one
contiguous
ifplus its comment), or just exportFLEET_LOAD_GUARD_ALLOW_SUPERSEDED=1. - PR #45: close it unmerged; it is comments only and touches no behaviour.
Outstanding
ISSUE-20260924-35 stays open at low severity: archive
the inert script + its unloaded plist to an
archive/<what>-<date>/ path. Move, never
delete. Deferred tonight only because it is launchd work.
Tickets
ISSUE-20260924-35opened, then corrected and downgraded (hazard neutralised).ISSUE-20260924-11/ISSUE-20260905-31confirmed genuinely resolved, fix verified on live data.- No leases held at end of session.
No secrets written.
Post-reboot continuation — appended 2026-09-24 23:16 EDT
Fleet rebooted 22:57:42 EDT (grok@rdmsm4x, on Rich's
behalf). This session was resumed by
resume_agent_sessions.zsh in a herdr tab. Every measurement
above predates that reboot; the two items below were measured after
it.
PR #45 merged
e5a663d Merge PR #45 — the
HostPressureLevel doc-comment fix is on Tyrell
main. Canonical tree clean, no orphan worktree registration
left behind.
A high-priority alert that had already self-resolved
ticket-store-commit: push rejected by backup (bus
20260924-223108, addressed to claude@rdmsm4x). Measured
rather than acted on: the job's own log shows
[PASS] pushed to backup (b960e1c) at 23:08:23 and
[PASS] pushed to fleet at 23:10:15, and
git rev-list --left-right --count HEAD...backup/main is
0 0. The 22:31 alert was true when sent and stale by the
time I read it. No merge or push was performed on the shared
ticket store — reconciling a divergence that no longer exists
would have been a write for nothing.
ISSUE-20260924-36 filed — the stability review false-fails hosts after every reboot
The 22:59 review graded rdmpw3275m FAILING
(0%) on a 60-second window at 0m uptime,
reporting "delivered 0 of ~1 expected runs".
Measured 17 minutes later:
| host | agentstatus | agentheal | review's verdict |
|---|---|---|---|
rdmpw3275m (uptime 17m) |
runs=3 interval 300 |
runs=2 interval 600 |
FAILING 0% |
rdmsm4x |
runs=3 |
runs=2 |
stable 100% |
Identical. Three runs of a 300s job in 17 minutes is exactly right; nothing was wrong with that host.
Root cause is the review's own rule not applied to its last case. Its
preamble says "A job cannot have run before it existed … grading an
unattended reboot against uptime produces a false FAILING grade",
and it applied that to rdmbair15m5 — then gave
rdmpw3275m a verdict instead. A 300s-interval job cannot
run once in 60s, so honest expected is 0; reporting "~1
expected" means the divisor is rounded up, which converts too early
to tell into a failing grade.
It matters because the fleet reboots all at once, so every host
enters that window together and the next post-reboot review will
false-fail several at a time. A review that cries failure after every
reboot trains its readers to discount it. Fix filed:
expected = floor(window/interval), allow 0, emit a distinct
"too early" verdict and exclude those hosts from the fleet
percentage.
Reply with the measurement: bus
20260924-231556-48AEACE9.
Publish status of this file
Written and correctly named for the autopublish job, which reports it
as new. The job has been running since boot and had not
reached it at 23:16. No second route was attempted: publishing via the
Notes MCP violates the llmlog standard and
poisons the job's duplicate check, and notes_changelog.zsh
hangs from a herdr pane on an unanswered TCC consent. The hourly job is
the correct and only route.