Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260916-1345-rtty-overnight-verified-and-session-handoff

rdmsm4x-changelog-20260916-1345-rtty-overnight-verified-and-session-handoff

2026-09-16 13:45 EDT · rdmsm4x · claude@rdmsm4x (Opus 5, session 90ecf89f)

RTTy Build 513's overnight soak passed, the p0 it was built for is resolved, and the remaining CPU-growth defect is now pinned to one specific question. Session documented for restart: nothing is in flight.

Companion to rdmsm4x-changelog-20260915-1950-rtty-soak-found-per-point-chart-recomputation.md, which covers the build and rollout.

What changed, and why it matters

Scope ~/dev/apps/RTTy (docs + tickets), ~/dev/issues (3 tickets), ~/dev/_handoff/rtty-diagnostics-20260916 (new)
Code changed? No. Documentation, tickets, and promoting measurement artifacts out of a session scratchpad.
Fleet state All six on Build 513, unchanged, 17 h uptime, zero diagnostic reports
Deployed? Nothing deployed today

1. Overnight soak passed — ISSUE-20260915-10 resolved (p0)

~13.7 h, then re-measured at 17 h. Every reachable host: one process, build 513, executable 5d5313dfd579, zero RTTy* diagnostic reports since the 20:45 rollout.

Fleet at 13:42 EDT — the last-3.3 h column is current behaviour, not the diluted average:

host uptime avg since launch last 3.3 h RSS
rdmpw3275m (Intel, 6K, dashboard open) 17.4 h 90.8 % 96.4 % 3076 MB
rdmsm4x 17.0 h 29.3 % 34.0 % 1315 MB
rdmbair15m5 17.0 h 13.0 % 1.5 % 150 MB
rdmbair13m5 17.0 h 3.6 % 2.7 % 126 MB
rdmpw3265m (no window) 17.0 h 3.5 % 3.5 % 1093 MB
jdmbair13m5 offline since ~22:00 — sleeping laptop. Took 513, verified 20:45:51.

2. The open defect is now one specific question (ISSUE-20260915-11, p1)

The finding that matters: the view tree does not grow. Same Mac, same window, footprint CoreAnimation regions — 309 at 7 minutes, 310 at 13.7 hours. One region in 13.7 hours. So the app is not accumulating views; SwiftUI's change-tracking is getting more expensive for a constant number of them.

At 14 h the profile is libswiftCore 47.3 %, libswiftObservation 24.4 %, libsystem_kernel 12.6 % (madvise the 5th hottest leaf), RTTy's own code 0.5 %. Optimising RTTy's own code cannot move this — which also disproves the "ECSNetTop per-sample payload" candidate I had named the night before.

Still not a leak, at the larger scale too: ps rss 2942 MB against footprint 343 MB dirty.

Workaround, shipped and proven over 17 hours: closing the dashboard window → ~3.5 % of a core.

3. Declined a shared-library adoption, with reasons (FEAT-20260830-45)

A peer lane's survey named RTTy as the one genuine consumer for a new ECS0SettingsStore, making RTTy the deciding vote on the epic's adoption gate. Declined, after reading all nine references: RTTyPortableSettings is a CloudKit wire payload that never touches disk — its schema version exists to reject an incompatible peer, not to migrate, and it has never had a migration. RTTy's real settings are UserDefaults (50 calls, load-bearing in qa_release.sh and every store test); replacing that would be a user-data migration with settings loss as the failure mode.

So the gate is met by zero genuine consumers, not one. Recorded on the ticket and sent back to the lane. Their library itself is not in question — additive, 31 tests, nothing calls it.

4. Measurement kit promoted out of the scratchpad

~/dev/_handoff/rtty-diagnostics-20260916/ (35 MB) — nine tools and six sample captures that cannot be regenerated, because the processes they came from are gone. Includes the before/after pair at matched uptime that proved the per-point fix. Read its README.md first.

Two method warnings are recorded there because each cost hours:

  1. sample leaf-share describes the composition of work, not its magnitude, and it pointed at two different libraries on two different hosts. Narrow with leaf ranking; conclude only from the call tree.
  2. Window visibility dominates CPU — ~3.5 % with no window vs ~96 % with a dashboard open on a 6K display. Check it before any cross-host comparison. The 513 rollout relaunched every app and window restoration reopened dashboards on hosts that previously had none, so rdmbair13m5's 1.7 % → 16.5 % was window state, not a regression.

Commands run (all read-only against the app)

ssh <host> 'zsh /tmp/soak_probe.zsh'        # ps-based; no mutation
ssh <host> '/usr/bin/sample <pid> 10 -file …'
ssh <host> '/usr/bin/footprint -p <pid>'
ticket claim / resolve / comment / reindex

No build was cut, nothing was deployed, no host configuration changed, no files deleted.

Verification

State changed / how to undo

Change Undo
ISSUE-20260915-10 resolved; -12 and -13 resolved yesterday ticket status <id> open
ISSUE-20260915-11 retitled twice to match evidence title is in the ticket's front matter
Comment on FEAT-20260830-45 (another project's epic) comments are append-only; add a correction
~/dev/_handoff/rtty-diagnostics-20260916/ created rm -rf that directory — but the samples are irreplaceable
2 commits on main (6fc67b7, a8a556f) git revert a8a556f 6fc67b7

Resume state

Nothing in flight. No ticket claims held, no background jobs running, clean tree, remotes in sync. The resume block is the top section of ~/dev/apps/RTTy/Project/SESSION-STATE.md.

Outstanding for Rich

Tickets

ISSUE-20260915-10 resolved · ISSUE-20260915-11 open, p1 · ISSUE-20260915-12 resolved · ISSUE-20260915-13 resolved · FEAT-20260830-45 decision recorded