rdmsm4x-changelog-20260915-1950-rtty-soak-found-per-point-chart-recomputation
2026-09-15 19:50 EDT · rdmsm4x ·
claude@rdmsm4x (Opus 5, session 90ecf89f)
Revised 21:00 EDT. The 19:50 version of this note said "fixed and committed; not yet released". That changed: the fix turned out to be a one-line hoist rather than a refactor, the A/B was decisive, and Build 513 went to all six hosts at 20:45. Everything below reflects the final state.
A three-hour soak of RTTy Build 512 confirmed the two fixes it shipped are holding, and found the real cause of the "graphs jump, parts disappear, it makes the system feel slow" report: the latency chart recomputed a whole-series cadence — including a sort — once for every point it drew, on the main thread inside the canvas display list. Fixed, A/B verified, and shipped.
What changed, and why it matters
| Scope | ~/dev/apps/RTTy (code + docs
+ tickets); Build 513 deployed to all six hosts |
| Released? | Yes. 0.3.838
Build 513, c71bc9b, tag v0.3.838-build513 |
| Fleet state | All six on Build 513, verified 20:45:51; rollback to 512 armed on every host |
The result, in one table
rdmpw3275m (Intel Mac Pro, 6K display, dashboard open) —
same host, same measurement script, the unfixed curve captured
before the fix was deployed there:
| uptime | Build 512 | Build 513 | |
|---|---|---|---|
| ~6 min | 33.3 % of one core | 33.0 % | 1 % |
| ~11 min | 43.1 % | 36.5 % | 15 % |
| ~16 min | 57.3 % | 41.8 % | 27 % |
| ~21 min | 64.5 % | 45.8 % | 29 % |
CORRECTED 22:30 EDT, after the full two-hour curve. The table above is the first 21 minutes, where the fix looks best. The advantage peaks there and then narrows: −22 / −19 / −15 / −13 % at 26 / 31 / 36 / 41 minutes.
Build 513 first reaches 73.6 % — the level 512 hit at 41 minutes — at 71 minutes. The fix buys about 31 minutes of delay, not a lower ceiling. From ~76 min it oscillates 76–83 % and sits at 76.1 % at 2 h 01 m; unfixed 512 was 82.6 % at 2 h 58 m. The plateaus are in the same band. RSS went 519 → 1570 MB over the two hours, essentially unchanged by the fix.
So the defect was real — the profile proves the sort is gone and RTTy's own code fell from 50.2 % to 3.3 % of working samples — but it does not solve the steady state. The growth ticket is not "largely fixed" and has been retitled to say so.
A hypothesis, not a conclusion: a plateau at ~80 % of one core with the work on a single main thread is what main-thread saturation looks like. If so the plateau is a ceiling rather than stabilisation, and redraw rate keeps degrading past it while CPU stops rising — which would match "feels slow". CPU alone cannot tell those apart; measuring redraw rate at 20 min versus 120 min would, and that is the next step.
1. Build 512 soak result — the 510–512 fixes hold
~3 h after the 16:02 rollout, all six hosts: zero crashes,
zero RTTy* DiagnosticReports, one process each
from /Applications/RTTy.app, build 512.
The hidden-dashboard fix is confirmed over hours rather than by spot
check — the three hosts with no visible window averaged 1.7 % /
3.1 % / 3.5 % of one core across the whole soak.
jdmbair13m5, which burned 149 % on Build
509, is at 3.5 %.
2. Found: CPU grows with uptime; a restart resets it (ISSUE-20260915-11)
rdmpw3275m (Intel Mac Pro, 6K display, dashboard open)
went 36.4 % of one core at 6.5 min uptime →
82.6 % at 2 h 58 m → 91–98 % at 3 h 17
m–3 h 25 m, still climbing.
A controlled restart on rdmsm4x, identical window
geometry and visibility, with the prediction registered beforehand
(10–18 %): 30.4 % at 3 h 04 m → 15.8 % at 2 m 45 s.
Ruled out with evidence, each: memory leak (ps rss reads
1.6 GB but footprint is 395 MB — the rest
is reclaimable malloc pages, i.e. churn, not retention), thread leak
(flat), unbounded event log (capped), write-during-read invalidation
(all @ObservationIgnored).
3. Found and FIXED: per-point recomputation (ISSUE-20260915-13) — the real cause
PathHistoryView.maximumLatencyConnectionInterval is a
SwiftUI computed property, so it is re-evaluated on
every access. It filters the whole sample array and then
sorts to take a median cadence.
shouldConnectLatency reads it, and that is called
once per adjacent point pair inside the plot loops. So
one canvas redraw cost O(points × history) where O(history) once was
correct — and the value is identical for every point in a single
draw.
It runs on the main thread inside the canvas display list, so it degrades redraw rate directly. That is the reported symptom.
Verified by profile, same host, same uptime, dashboard open:
| Build 512 @ 13 m 40 s | fixed @ 13 m 35 s | |
|---|---|---|
| working (non-idle) leaf samples | 2,211 | 1,268 |
| RTTy's own code | 1,111 (50.2 %) | 113 (8.9 %) |
_stableSortImpl |
501 | gone |
_merge |
240 | gone |
observedCadence(in:within:) |
168 | gone |
sort() |
95 | gone |
The fix is pure hoisting onto the view's existing render memo, whose key already covers every input of the computation. Identical output; only the number of times it is computed changes.
An earlier estimate of mine was wrong and is recorded as such: a benchmark put this function at ~0.3 % of one core, assuming a ~1 Hz call rate. The profiler showed the real call site is per drawn point. Measure the call rate; do not assume it.
4. Also fixed: menu-bar cadence measured across all history (ISSUE-20260915-12)
menuBarBucketDuration measured the cadence binding a
5-minute grid across all retained
history (up to 360 min/target), so it lagged cadence changes and its
median-sort grew O(n log n) in uptime: 0.020 → 0.536 → 1.125 ms at 390 /
10,800 / 21,600 samples per target, against a flat 0.025
ms once windowed. Added windowBoundCadence rather
than changing existing semantics; the window is a floor, not a
cap, so a slow probe cadence is still measured.
Commands run (all reversible)
# measurement only
ssh <host> 'ps -axo pid=,etime=,time=; /usr/bin/sample <pid> 10; /usr/bin/footprint -p <pid>'
screencapture -l <CGWindowID> # window-scoped, never -R
# app restarts (reversible; each came back with the same window geometry)
osascript -e 'tell application "RTTy" to quit'; open -a /Applications/RTTy.app # rdmsm4x, rdmpw3275m
# release build, gates, notarize, deploy
RTTY_CODESIGN_IDENTITY=<devid-sha1> RTTY_PROVISIONING_PROFILE="<developer-id profile>" \
bash script/qa_release.sh # 10/10, rc=0
RTTY_NOTARIZE=1 ... ./script/build_and_run.sh --stage # notarize + staple
zsh ~/dev/_handoff/rtty-build513-20260915/release/build513/tools/deploy_build513_fleet.zsh <host>Verification
- Full suite 1060 / 71 / 285, 0 failures (was 1050/71/285 before tonight; +10 new tests), and the release gates ran the same counts natively and under Rosetta.
- Every new test was confirmed failing before its fix. The 5 cadence tests each failed with the exact value predicted from the quantum ladder beforehand (7.5, 15.0, 3.75).
- Both fixes mutation-tested. The memo: removing it
entirely fails 3 tests (900 computations vs 1); treating a cached
nilas a miss fails only the test written for that case. The cadence sparse-window floor: removing the clamp fails only its own test, 3.75 vs 15.0. - Remote heads independently verified at
9ad1003on bothfleetandbackup, tagv0.3.838-build513present on both. - Fleet state verified by a separate script run on each host, not from the deploy's own report.
State changed on this host / how to undo
| Change | Undo |
|---|---|
| Build 513 installed on all six hosts | /Applications/.RTTy.app.rollback-build512-c71bc9b
is kept beside the app on every host — swap it back to return to Build
512 |
RTTy restarted on rdmsm4x and
rdmpw3275m for the controlled experiment |
Superseded; the 513 rollout replaced every process anyway |
dist/RTTy.app rebuilt from
HEAD |
dist/ is a build artifact and
gitignored |
6 commits + 1 tag on
main |
git revert 9ad1003 c71bc9b 82dc531 0233e56 bc27473 0ef7200 bcd8dfa;
git tag -d v0.3.838-build513 |
No files were deleted and no host configuration was changed. The only host change is the app itself, which is reversible per-host from the rollback copy.
Build 513 release facts
Gates 10/10, rc=0 (20:12–20:24 EDT). Native and
Rosetta both 1060 (8 skipped) / 71 /
285, 0 failures; Intel suite 1408 of 1416 discovered cases.
Notarized — submission
c63486ac-a535-4b25-9325-8b5445db55e9,
Accepted — stapled, spctl = Notarized
Developer ID, lipo = x86_64 arm64, embedded
profile 0efc0851… unchanged from 512.
Rolled out canary-first: rdmpw3275m 20:21:49, then
rdmbair15m5, rdmbair13m5,
jdmbair13m5, rdmpw3265m, hub
rdmsm4x last, 20:39:42–20:45:25, each accepted in 54–60 s.
Independently verified 20:45:51 (not trusting the
deploy's own report): all six on build 513, executable
5d5313dfd579, one process each from
/Applications/RTTy.app, zero diagnostic reports.
Outstanding
- ISSUE-20260915-10 stays open until the overnight fleet check. The soak clock restarted at 20:45 with Build 513.
- ISSUE-20260915-11 open, and not largely fixed — see the correction above. Build 513 still reaches ~80 % of a core within ~75 minutes on the Intel host.
- The next candidate is NOT the
ECSNetToppayload, though an earlier draft of this note said it was. A post-fix sample at 29 minutes puts RTTy's own code at 3.3 % of working samples, so no optimisation of RTTy's own sorts can account for the remainder. What is left is SwiftUI itself (libswiftCore 37.1 %, libswiftObservation 12.0 %, malloc 11.9 %, AttributeGraph 9.6 %). - Fleet caveat: the rollout relaunched every app and
window restoration reopened dashboards on hosts that
had none.
rdmbair13m51.7 % (no window, 512) → 16.5 % (window on screen, 513) is window state, not a regression.rdmpw3265mhas no window and is 3.1 % on both builds, so the hidden-dashboard gate still holds. Control for window visibility before comparing hosts. jdmbair13m5took 513 at 20:43 and was verified at 20:45:51, then went off the network (100 % packet loss) — a sleeping laptop. Re-check in the morning.- App Store Connect still holds Build 509, which lacks every fix from 510 onward. Attaching a newer build is Rich's call.
Tickets
ISSUE-20260915-11 (p1, open) ·
ISSUE-20260915-12 (p2, resolved) ·
ISSUE-20260915-13 (p1, resolved)
Evidence
~/dev/apps/RTTy/docs/release-evidence/build-512/soak-2026-09-15.md
· ~/dev/apps/RTTy/Project/SESSION-STATE.md