rdmbair15m5 is back on the tailnet: tailscale up brought
its already-authenticated node from Stopped to
Running, and the scheduled (launchd) fleet probe now sees
6/6 hosts, verified via a forced launchd kickstart, not just an
interactive run.
Scope
- rdmbair15m5 — state change (Tailscale daemon
brought up). Reached via
[email protected](LAN). - rdmsm4x — local file edits only (this is the host this session ran on).
- No other fleet host touched.
What was wrong
tailscale statuson rdmbair15m5 (via/Applications/Tailscale.app/Contents/MacOS/Tailscale status --json) showedBackendState: Stopped,HaveNodeKey: true,AuthURL: ""— the node was already registered/authenticated, just not connected. Not a login problem.- The
tailscaleCLI was already present at/opt/homebrew/bin/tailscale(Homebrew formula, installed 2026-08-20 21:41), just not on the restricted PATH of a bare non-interactive SSH command (ssh host 'cmd'uses/usr/bin:/bin:/usr/sbin:/sbinonly). A login shell (zsh -lc, which is how launchd invokes scripts) already resolved it fine. ~/dev/data/fleet_hosts.jsonhad rdmbair15m5's two tailnet addresses in the wrong order: the stale duplicate node's IP (100.75.253.27) came before the live node's IP (100.74.59.4) in the array, sofleet_status.zsh'scandidates()(rank-0-with-stable-tiebreak) would try the dead address first on every future probe.
Commands run
(on rdmbair15m5, over [email protected])
brew install tailscale # already installed/up to date (1.102.3) — confirms CLI presence
tailscale up # BackendState Stopped -> Running, no re-auth prompt (rc=0)
tailscale status --json # confirmed: Running, Self.Online true, direct connection
Commands run (on rdmsm4x)
python3 <one-off script> # reordered rdmbair15m5's entry in data/fleet_hosts.json using
# live `tailscale status` as ground truth; every other host's
# cache was already correctly ordered (no-op for them)
touch -t 202608220001 ~/dev/data/fleet.json
launchctl kickstart -k gui/501/com.eastcoastscience.devsite
Verification evidence
ping 100.74.59.4from rdmsm4x: 3/3 packets, 0% loss.ssh [email protected]from rdmsm4x: connects, returns hostnamerdmbair15m5.tailscale pingfrom rdmbair15m5 to rdmsm4x: direct (non-relayed) connection.~/dev/scripts/fleet_status.zshrun interactively: 6/6 up.- The path that actually matters: after forcing
data/fleet.jsonstale and kickstartingcom.eastcoastscience.devsite(the real launchd job, not an interactive shell),data/fleet.jsonshows all six hostsup:true, anddata/publish.logrecorded2026-08-22 19:15:45 EDT ok 10 product pages · 201 events · 6/6 hosts upfor that exact forced run (not a cached "(local only)" line).
Files touched
~/dev/data/fleet_hosts.json— reordered rdmbair15m5's tailnet addresses (live IP first). Diff was a two-line swap, nothing else changed.~/dev/data/status-updates/fleet.json— created, shape{"name":"Fleet","tested":[...],"untested":[...],"nxt":[...],"blocked":[...]}.~/dev/RESUME.md— appended a RESOLVED entry under the existing "OPEN — rdmbair15m5 is off the tailnet" heading.~/dev/data/fleet.json— rewritten by the launchd job itself (not hand-edited), now shows 6/6.
Backup / undo
- Pre-edit copy of
fleet_hosts.jsonis at/private/tmp/claude-501/-Users-richh-dev/475a6c9b-a641-4e67-9688-2eb09ec4b07c/scratchpad/fleet_hosts.json.before(session scratchpad — not durable; the change itself is a 2-line array reorder, trivially reversible by hand if ever needed). tailscale upon rdmbair15m5 has no meaningful "undo" — it just resumed the existing, already-authorized connection.tailscale downwould put it back toStoppedif ever wanted.
Outstanding — needs Rich, not blocking
Every fleet host (rdmbair15m5, rdmbair13m5, rdmpw3265m, rdmpw3275m,
and rdmsm4x itself) is carrying one stale duplicate Tailscale
node: a live <host>-1 entry plus a bare
<host> entry that's been offline 1+ day. Harmless
(the live node is what's used) but cluttered, and only the account owner
can remove another node's registration — the CLI can't. To clean up:
https://login.tailscale.com/admin/machines
→ for each of the 5 stale (non--1, "Last seen" 1d+) rows →
⋯ menu → Delete. Double-check you're
deleting the offline row, not the -1 (active) row, before
confirming.