rdmsm4x-changelog-20260916-1103-audit-dedupe-fix-ticket-consolidation-estate-push
Run 2026-09-16 10:25:53 → 11:03:37 EDT on
rdmsm4x by claude@rdmsm4x/dev-c5 (cli
session 65c5e356). All timestamps read from
date. Budget proceed on both vendors
(Anthropic 7d 31%, OpenAI 7d 7%); no subagents spawned, no external
spend.
One line: found and fixed the reason the fleet ticket store was filling up — the nightly audit's dedupe lookup had never once matched, filing a fresh ticket every night — then consolidated the 91 duplicates it had produced, and levelled the git estate with 38 fast-forward pushes and 5 fast-forward pulls.
| Measure | Before | After |
|---|---|---|
| Open tickets | 439 | 345 |
Open nightly-audit tickets |
103 (for 12 real conditions) | 11 |
| Resolved tickets | 531 | 632 |
| Repos not level with a remote | 30 ahead / 8 behind | 3, all diverged or vendor forks |
| Open PRs across 362 GitHub repos | 1 (draft, CI red) | 1 — correctly left alone |
1. Root cause: the audit's dedupe had never fired
fleet/ops/audit/nightly_audit.zsh looked up its existing
ticket with
existing=$(ticket search "nightly-audit $cid" | grep -oE "^[A-Z]+-[0-9]{8}-[0-9]+" | head -1)
but ticket search renders each hit as
• ISSUE-20260913-08 [issue|open|medium] (fleet): …. The
line starts with a bullet, so the ^ anchor
matched zero lines, $existing was always
the empty string, the else branch ran, and a brand-new
ticket was filed every night.
The block's own comment describes a careful content-aware dedupe — subject-set fingerprints, "changed subjects → comment AND say what changed". None of it was ever reached.
Why it survived two weeks undetected: a failed match and "there is no existing ticket" are the same empty string. The broken path is indistinguishable from the success path, so nothing downstream could notice. This is the recurring fleet failure mode — a check that measures the wrong thing and returns a confident wrong answer.
Scale by 2026-09-16: 103 open tickets for
12 distinct conditions — intel-bottles 14,
repos-no-remote 12,
repos-unpushed/dead-sessions/bus-unread-high/
boot-contract 11 each, version-cadence 10,
lanes-reconcile 9, spoke-signing 7,
mem0-fleet 5, silent-write-back 1,
prs-open-24h 1.
The fix — fleet@e4513be
Parses ticket search --json instead of scraping the
rendered output, and tightens two things the old code got wrong even in
principle:
status == "open"only. The old lookup would happily comment on a resolved ticket, where the comment is silently lost.- Exact check-id title prefix, so a substring hit on a different check cannot be selected.
- Lowest id wins, so the surviving ticket is stable run over run rather than whatever the search happens to rank first.
Verified in both directions (the rule is that a fix must not overshoot): all 12 live check ids resolve to their oldest open ticket, and two non-existent check ids return empty.
boot-contract -> ISSUE-20260906-11 intel-bottles -> ISSUE-20260902-12
bus-unread-high -> ISSUE-20260906-06 repos-no-remote -> ISSUE-20260905-03 (…12 total)
this-check-does-not-exist -> [] zzz-nonsense -> []
2. Consolidation — 91 duplicates resolved into 12 survivors
For each check id: the lowest open id was kept (the
same ticket the fixed lookup now selects, so tonight's run lands on it),
a consolidation comment was appended to it carrying the most
recent observed state forward so nothing was lost, and every
other ticket was resolved with --reason naming the survivor
and --evidence citing e4513be.
90 of 91 resolved. The one refusal was correct and was left
alone: ISSUE-20260916-03 is claimed by
agy@rdmsm4x/standby (claimed 10:27, heartbeat current,
lease to 13:27). ticket resolve refused,
--force was not used — forcing past a live
holder is exactly the collision the lease exists to prevent. A comment
was left on it explaining the consolidation and leaving the decision to
agy.
Counts re-read from the store afterwards rather than trusted from the
log: open 439 → 345, open nightly-audit 103 → 11, resolved
531 → 632.
3. Pull requests — 1 open, correctly not merged
362 GitHub repos swept with a per-repo gh pr list (the
gh search prs index lags and was not trusted). One open PR
fleet-wide:
rdmsm4x-dev-apps-tyrell #35 — "Bound
pending upload retries to missing inventory hashes", branch
codex/tyrell-build21-memory-20260914. It is a
draft, and its
swift build && swift test check is
FAILURE. Not merged, not modified: a draft with a red
build is not mergeable, and it is another lane's in-flight work. It is
already tracked by the audit's own prs-open-24h ticket
(ISSUE-20260916-07).
4. Git estate — 38 pushes, 5 pulls, 0 failures
Same discipline as always: plain git push (never
--force, so a non-fast-forward is rejected rather than
overwriting), merge --ff-only on clean trees only, and
remotes chosen by inclusion — URL under
github.com/richhdoty/ or git.ecs0.net — which
automatically excluded every
legacy/DO-NOT-PUSH/upstream
remote and both vendor forks.
Pushed (38): bookmarkROO, cvedb, dataROO, devSort,
DishTTY, fileLabeler, hyperFile, iCloudMonitor, LogTTY,
macOS-ramDisk-app, PasswordScope, rdDB, rdKIPP, RDnamed, ReceiptRoo,
rooDB, RTTy (origin +46, then +1 more as its lane committed mid-run),
scanRoo, SQLiteScope, statement-collector, SubnetCalculator, Xentropy,
xTTY, design, net/fleet-netmgmt, sites/dev.ecs0.net, fleet — across
backup, fleet and origin as each
lagged.
Fast-forwarded (5): CoreAudit +3, ModelVault +4,
ThermalScope +2, ramdisk +6, SQLiteScope +2 — each proved a true
fast-forward with merge-base --is-ancestor first, then
re-pushed so all three mirrors ended level.
Secret-scanned before publishing. Two hits, both
false positives: a sentinel string
let token = "POSIX_SETTLE_BARRIER_OK", and prose reading "a
synthetic secrets file" in a test description.
Verified by an independent post-push sweep of all
183 checkouts, not by trusting push output. Three repos remain, all
needing an owner's merge decision, none automatable:
apps/macOS-ramDisk-app_eval (diverged 20/1 on all three
remotes) and the upstream vendor forks lib/t3code (926
behind) and lib/tokenbar (188 behind).
5. Two survivor conditions re-measured — neither resolved
Measured with the audit's own check (sourced
fleet/ops/audit/lib/repository.zsh and ran CHECK 2's loop
verbatim, positive control: 176 repos enumerated, not 0).
| Check | Was | Now | Verdict |
|---|---|---|---|
repos-unpushed |
4 | 2 | Both survivors are the vendor forks lib/t3code /
lib/tokenbar, whose commits can never go to origin. Needs a
mirror, not a push — outward-facing, so Rich's. |
repos-no-remote |
3 | 5 | Worsened. net/fleet-certs,
lib, apps/xReaper plus new
sites/firmwaredb.app, apps/xcpe. Invisible to
the backup pipeline. |
Both got a comment with the measurement; neither was resolved, because neither count is zero.
A wrong measurement I made, and how it was caught
My first attempt at this ran the check without sourcing
lib/repository.zsh. In zsh,
! <command-not-found> evaluates
true, so every repo counted as unpushed — 171
of 176, reported without an error. The negative control caught
it: apps/RTTy was in the list, and I had just verified it
level on all three remotes. Sourcing the library gave 2. The lesson is
the standing one — a probe needs a control that would fail if the probe
were broken, and "the number looks plausible" is not that control.
6. A shared-tree hazard worth recording
Running nightly_audit.zsh at 10:33 produced
parse error near 'fi' at line 520 in a
file that was 516 lines long. That impossible line
number was the tell: another agent was rewriting the file at that moment
(their commit 990275d landed with mtime 10:36:56), and the
run read a torn file. Nothing was wrong with either edit — both commits
are in history and zsh -n passes. When running a
script out of a shared tree, run a snapshot copy, which is what
the successful re-run did.
7. Outstanding — owner actions
- Mirrors for
lib/t3codeandlib/tokenbar— local commits exist on no backup at all. - Remotes for the 5 no-remote repos —
git.ecs0.netis the lower-risk destination; none is obviously publishable. Also:~/dev/libitself is now a git repo, which is probably unintended and worth a look before it is given a remote. apps/macOS-ramDisk-app_evaldiverged on all three remotes — needs a merge decision.- Tyrell PR #35 — draft with a red
swift build && swift test; its author's to fix. ISSUE-20260916-03— left withagy@rdmsm4x/standby, who holds the lease.
8. How to undo
- The audit fix:
git -C ~/dev/fleet revert e4513be. A pre-edit copy is at/tmp/nightly_audit.zsh.bak-20260916-102753. - The 90 resolved tickets: each carries a
--reasonnaming its survivor and a claim/resolve history; the ticket engine keeps them inissues/resolved/and they can be reopened. The full per-ticket log is at…/scratchpad/consolidate-20260916.log(90rc=0, 1rc=3). - The pushes: every one was a fast-forward that only added commits. The old SHAs are in §4's ranges and in each remote's reflog.
- The pulls:
git -C ~/dev/apps/<name> reset --hard <old-sha>— CoreAudit00e775b, ModelVault1a8c450, ThermalScope3f34bf3, ramdisk2ba0c3f, SQLiteScopec17866a.
9. Evidence
Scratchpad
/private/tmp/claude-501/-Users-richh-dev/65c5e356-…/scratchpad/:
pr-sweep-20260916.txt ·
repo-sweep-20260916.txt (before) ·
repo-sweep-verify-20260916.txt (after) ·
consolidate-20260916.log.
No secret value appears in this changelog. Credentials are referenced by name and location only.