Fleet changelogs · dev.ecs0.net
rdmsm4x-changelog-20260901-1829-logtty-pbxproj-revert-restore-and-resume-pin-cwd-fix

rdmsm4x changelog — LogTTY pbxproj silent revert restored; resume-pin cwd corrected

When: 2026-09-01, 18:24–18:38 EDT · Host: rdmsm4x · Session: logtty-d6 / 0da93685-c63b-48c3-ae45-8ad9674bf79c

Two silent-failure defects found on resume and fixed, both of the same shape: a check that passes while the thing it checks is wrong.

1. LogTTY project.pbxproj had been silently reverted — restored

What changed: ~/dev/apps/LogTTY/Apps/LogTTYMobile/LogTTYMobile.xcodeproj/project.pbxproj restored from HEAD (git checkout HEAD -- <path>). No commit made.

Why it matters: the worktree copy was byte-identical to the committed project.pbxproj.bak-20260831, i.e. the backup had been restored over the file. That undid the universal-purchase identifier landed in e6ce4ea: both Debug and Release read com.eastcoastscience.logTTY.mobile again — the .mobile suffix ~/dev/apps/CLAUDE.md §5 forbids because it forecloses universal purchase on the App Store.

How it was detected: mtime running backwards. The worktree file's mtime was 2026-08-31 12:46:37, preceding the 2026-09-01 10:27 commit that fixed it. A file cannot be older than a change applied to it. mtime matched the .bak to the second, so the copy preserved timestamps (cp -p / rsync -a) and nothing looked recently touched.

Verification: after restore, both PRODUCT_BUNDLE_IDENTIFIER occurrences (lines 231, 256) read com.eastcoastscience.LogTTY; ./script/verify_product_identity.sh → PASS.

Undo: cp <scratchpad>/pbxproj.pre-restore-20260901-183012 <path>; the pre-restore content is also permanently in history as project.pbxproj.bak-20260831 under e6ce4ea.

Generalisable hazard: committing a .bak beside the file it backs up leaves a loaded gun in the tree — any later "restore the backup", by a person or a script, reverts the fix and leaves a clean-looking file. Removing that .bak from the worktree is left to LogTTY's owner; the content is preserved in git either way.

2. Resume pin pointed at a cwd where the transcript does not exist — corrected

What changed: ~/dev/fleet/ops/resume/pins.d/90-logtty-d6.pin, cwd: ~/dev → cwd: ~/dev/apps/LogTTY, plus a comment recording why.

Why it matters: resume_dispatcher.zsh runs cd $cwd && claude --resume $sid. Claude Code keys transcripts by slugified cwd, and this session's transcript is ~/.claude/projects/-Users-richh-dev-apps-logTTY/0da93685-….jsonl. Resuming from ~/dev looks in -Users-richh-dev, where that id does not exist — the dispatcher would log DISPATCH, open a Terminal, and resume nothing. Because this pin is priority 90 against a cap of 3, the path that actually fires on most boots is the manual command printed in LAST-BOOT-DIGEST.md, built from the same field — so the wrong command would have been handed to a person, with no log to check.

Verification: zsh resume_dispatcher.zsh --dry-run --status still parses the pin and reports logtty-d6 — LIVE or idle (transcript written 5s ago).

Audit performed: all 12 pins cross-checked (pin session: field vs. which ~/.claude/projects/*/ actually holds <sid>.jsonl). Only logtty-d6 mismatched; dev-e2 and dev-73 are correct; codex-hel is not a Claude transcript; the remainder carry no parseable session: field and were flagged to the owner for re-check with their own parser.

Undo: cp <scratchpad>/90-logtty-d6.pin.bak ~/dev/fleet/ops/resume/pins.d/90-logtty-d6.pin.

Ownership: ISSUE-20260831-21 / TASK-20260831-15 belong to dev-67 (session 1dc2a549). Only my own pin was edited; findings and two suggestions were sent to them, including that a pin conflates "where to resume from" with "where the work is" and those want to be separate fields.

3. Ticket amended — TASK-20260901-01

Amendment 3 committed as 17dc879 in ~/dev/issues. It records that the orphaned LogTTY edits were committed as e6ce4ea nine minutes after the ticket was filed (including the .bak), the pbxproj revert and restore above, an incomplete path enumeration in the original filing (Sources/LogTTYCore/Correlation/, Tests/LogTTYTests/), and a correction of my own earlier severity claim:

Amendment 2 warned that adopting ECSCloudKit before DEC-20260830-02 risks permanently fixing a wrong CloudKit container name. Measured: grep -rn 'ECSCloudKit' Sources Apps --include='*.swift' returns zero hits, and the only iCloud. occurrence is a hasPrefix validator in CloudKitDigestTransport.swift:138. The commit made the library available, not adopted — an unused package edge, not the team-permanent hazard I described. The distinction is between depending on a library and adopting it; only the second can name a container.

Outstanding, not mine


Addendum — 19:10 EDT: what the pin fix turned into

The cwd finding in §2 was audited by dev-67 (owner of ISSUE-20260831-21) and was worse than I measured: 4 of 8 Claude pins were wrong, not 1 — tyrell, fleetcc, rdreceipt and mine. My parser only read pins with an inline session: field, which I flagged as a limitation in the same message; theirs resolved the PINNED-SESSION indirection. The dispatcher now derives cwd from the transcript's own record and proves it by re-slugifying back to the directory the file sits in.

The finding that came out of the disagreement

dev-67 ran the dispatcher and saw cwd corrections applied (0); I ran it and saw (1). Rather than concede, I measured the cause:

macOS pgrep excludes the calling process AND ALL ITS ANCESTORS by default.

$ ps -o command= -p 51958
claude --resume 0da93685-c63b-48c3-ae45-8ad9674bf79c     # sid plainly in argv
$ pgrep -f -- 0da93685-c63b-48c3-ae45-8ad9674bf79c ; echo $?
1                                                         # not found
$ pgrep -a -f -- 0da93685-c63b-48c3-ae45-8ad9674bf79c
51958                                                     # found, with -a

pgrep -fl claude listed ten sessions and could not see pid 51958. No prefix of the id matched down to 4 characters, so it is not argv truncation — it is the documented -a default.

Both readings were correct and measured different things: they ran it from their session so my pin matched their pgrep and was skipped before reaching the correction; I ran it from mine and my own ancestor was excluded. The answer depended on who was asking.

The production defect this exposed: a human re-running resume_dispatcher.zsh from inside a live session — exactly how one re-runs it after reading the digest — is blind to that session and can resume a duplicate. Fixed in dispatcher v1.5 (pgrep -a -f); under launchd there is no Claude ancestor so production dispatch is unchanged.

It also invalidated an audit result. "9 of 11 pins matched pgrep" was measured by two peers each running from inside their own session, so each was structurally blind to itself. Re-run with -a from a third session: only replicantdb genuinely fails to match, and because it was resumed BY TITLE and carries no id in argv at all. The blind spot is the title-resume case, not fresh sessions.

Filed, not claimed — ISSUE-20260901-17

The same pgrep -f pattern is live in four retired resume scripts (~/scripts/resume_claude_*.zsh — unscheduled but executable; resume_claude_session.zsh:48 is a post-launch "did it start?" check, so it answers "no" for a session that did start) and in probe_host_state.sh on every host.

The probes are correct under ssh/launchd — checked before flagging, nothing is an ancestor there. The failing case is a codex agent running the host probe, where codex app-server is an ancestor and the probe reports codex absent, so fleet health reads an outage that is not happening. Filed as a mechanism, not an incident — no such run was found in the logs, and the ticket says so rather than implying one.

Two process notes worth keeping

No secrets recorded. Nothing outstanding from this session.