Fleet changelogs · dev.ecs0.net
rdmsm4x - ops - Efficiency source fixes - codex - TASK2026092808 - fleet - 20260928-1147

Ops efficiency followup result

2026-09-28 11:45:19 EDT · rdmsm4x · ops@rdmsm4x/opseff0928

Work began 2026-09-28 11:25:22 EDT. TASK-20260928-08; child TASK-20260928-09; rollout decision DEC-20260928-03. Rich explicitly authorized integration, source fixes, controller non-objection, tests, rollback and dry-run measurement. No human messages, purchases, credential changes, irreversible deletion, or session termination. Usage gate returned conserve; no workers spawned. Heavy verification ran at nice 10.

Delivered and integrated

Status producers and deployment gap

The 1,785 baseline messages are seven known machine-status classes; their bodies were not all identical. TASK-20260927-72 had already migrated canonical producers to telemetry. This followup found all five installed spoke coordinators still using inbox mail. Only their bus-call argument list was backported; unrelated v1.15 behavior was preserved.

Class Baseline Canonical producer Result
fleetstate-* 926 ~/scripts/fleet_state_report.zsh Already telemetry on six hosts; verified
agent-coordinator-* 428 fleet/agent-coordinator/agent_coordinator.py; installed under ~/.agent-coordination/coordinator/bin/ Hub telemetry verified; five spokes repaired
devmon-daemon-restart* 125 dev/_ops/devmon/devmon.sh Already telemetry on six hosts; verified
agent-memory-advisory* 123 same devmon producer Already telemetry on six hosts; verified
uncommitted-work-in-* 106 fleet/maintenance/scripts/shared_tree_guard.zsh Hub telemetry producer verified
fleet-audit* 59 dev/scripts/fleet_audit.zsh Hub telemetry producer verified
devmon-daemon-alert* 18 same devmon producer Already telemetry on six hosts; verified

All six installed buses now have exact-body append dedupe. Latest snapshots retain a fresh timestamp on every observation; changed values and recoveries still append. No status/PID/number normalization hides changed alerts. Existing telemetry filenames, frontmatter and replication remain intact; real agent-to-agent mail is unchanged. Additional machine-topic scan found actionable ticket-store push/security failures and CVE alerts, which were retained; stability-review already uses telemetry.

Dry-run measurement Before After
Historical seven-class status messages entering inbox 1,785 0
Latest telemetry snapshots from that replay — 36
Historical append records 1,785 1,785
Identical-body fixture append records 2 1
Installed spoke coordinators still sending mail 5 0

The replay used actual agent_msg.zsh in an isolated HOME/coordination root with sync disabled. All 1,785 historical bodies differed, so no historical log-volume reduction is claimed from exact-body dedupe. The mailbox reduction includes the earlier migration; this task completed its missing deployment. This is not an observed future-week saving. Six remote hash checks and six isolated producer-call smoke checks passed. We did not run full coordinator jobs, restart agents, or claim a naturally scheduled pass as proof.

Evidence: inventory before, inventory after, replay, rollout, spoke rollback, fleet verification. Source replay script is in the canonical kit.

Abandoned claim causes and source fixes

The baseline has 81 force-release events; 73 explicitly describe stale/dead/expired claims. This is 70 distinct tickets, not 73 active sessions. Among those 73, former holders were Claude 37, Codex 5, Tyrell 4, xattr 3, launch 3, domains 3, rdmactccfix 2, library 2, and 14 other single-event lanes. Largest individual holders across all 81: claude@rdmsm4x/tyrlead1 12 and claude@rdmsm4x/L-ops 8.

The top recorded causes are session exit/lapse with no heartbeat or closeout: 36 events in the Sep 27 stale/dead batch and 13 in the Sep 28 lapsed/no-heartbeat batch. Four more cite >12h without a heartbeat; four XEntropy releases cite post-reboot process/worktree evidence. Smaller cases name completed headless workers, a failed driver, blocked owner operations, and handoff/closeout omissions. Bot lane names do not establish their underlying harness. Historical release reasons cannot distinguish every true exit from false staleness. See claim analysis for per-event identity/reason evidence.

Two source defects were addressed:

  1. Renewal lost on reindex. _extend updated only derived SQLite. A comment/reindex restored the acquisition expiry from Markdown. Three regression cases failed before; all now pass. Heartbeat snapshots persist/coalesce; original acquisition and native provenance survive rebuilds. Failed persistence fails renewal instead of reporting success.
  2. No reliable exit closeout. Hub Claude had no SessionEnd claim cleanup, and named ticket sessions hid native transcript provenance. Acquisition now records an exact, unambiguous native filename through metadata-only lookup; no native contents are read. The installed SessionEnd hook releases only that exact transcript on that exact host. --if-transcript rechecks provenance under the ticket write lock. No --force, no Stop cleanup, no parent/subagent cleanup, no task resolution, no process control.
Claim verification Before After
Heartbeat persistence regressions 3 failing / 5 cases 0 failing; expanded to 8 passing cases
Isolated terminating-session claim 1 0
Same claim after ordinary Stop 1 1
Task status after automatic release in-progress in-progress

The cleanup targets the largest measured harness group, hub only. New Claude sessions load the installed hook; existing sessions were not restarted. Natural session-end capture has not yet been observed. Missing/ambiguous provenance is deliberately skipped; crashes, SIGKILL and power loss still fall back to bounded leases. No claim of 73 prevented events or measured next-week improvement is made.

Verification totals and rollback

232 passing automated checks: 130 bus, 69 existing claim/lease, 8 new renewal/provenance, 10 cleanup/installer, 6 telemetry, 9 census. Also: six host hash checks, six isolated producer smokes, and the installed-hook isolated lifecycle scenario. Original bus suite exposed a pre-existing test leak into the live localdb reader and assumed host ordering; fixture isolation and explicit host selection fixed it. Final suite is 130/130, not a waived failure.

One-step fleet telemetry and hub-hook rollback (preserves backups and refuses later drift):

zsh /Users/richh/dev/fleet/maintenance/scripts/rollback_efficiency.zsh --apply

Individual host rollback: python3 ~/dev/fleet/maintenance/scripts/deploy_efficiency.py --host HOST --rollback. Hub hook only: python3 ~/dev/fleet/maintenance/scripts/install_session_claim_cleanup.py --rollback. Hub and spoke backup/rollback/reapply paths were exercised. Backups on each host: ~/.agent-coordination/backups/opseff0928/ with before/after SHA-256 manifest. The wrapper itself was syntax checked and dry-run verified; full simultaneous rollback was not run. Independent one-step claim rollback: zsh ~/dev/fleet/maintenance/scripts/rollback_claim_efficiency.zsh --apply. It removes only this hook and creates non-force Git reverts of a78dc8c and a5c224f, preserving ticket data and history. The wrapper refuses edited source paths, was syntax/dry-run checked, and has not been executed on canonical issues. Commits remain available for reapplication.

Limits and durable closeout

The requested old topology-audit path is absent; its canonical replacement at fleet/dev-fleet-reconciliation/scripts/audit_codex_fleet_topology_v1.0.zsh ran. Six hosts were reachable; three topology expectations failed (rdmpw3265m account-lane mismatch, jdmbair13m5 unclassified account lane and project-root state). No account/config/history was moved or altered to address those adjacent findings. Audit.

DEC-20260928-03 was sent to the controller before mutation; no objection received. Its rollout record was appended to SETTLED-ANSWERS.md. Existing unrelated dirty work retained. Apple Notes publication is pending: launchctl managername reports Background. The fleet-notes-publish rule requires a file copy in this case; it is saved in the changelog archive for the existing Notes autopublisher. No human intervention requested.

Closeout update: source race guard merged at maintenance 5da5612 and issues 04b9fbd; TASK-20260928-09 and DEC-20260928-03 resolved with evidence. Artifact tickets carry canonical integration receipts. Evidence/rollback record commit: 9fd3b29. Ticket closeout commit: 96e540a. Both main tips were read back from fleet and backup and matched exactly at closeout.

TASK-20260928-08 is resolved and both task leases are released. Notes tool returned exit3 (Background, not Aqua); file archive: /Users/richh/dev/LLM/Claude/changelogs/rdmsm4x - ops - Efficiency source fixes - codex - TASK2026092808 - fleet - 20260928-1147.md.

Final verification: zero claims for ops@rdmsm4x/opseff0928. Worktrees preserved under ~/dev/_worktrees/archive/ops-eff0928-maintenance and ops-eff0928-issues; no files deleted. Installed guarded SessionEnd lifecycle was rerun successfully after the race-guard merge. Internal completion receipt: bus 20260928-114839-8CCE30B8.