Fleet changelogs · dev.ecs0.net
jdmbair13m5-changelog-20260916-1641-tailscale-repair-and-fleet-naming

jdmbair13m5 changelog — Tailscale repair + fleet naming without /etc/hosts

Session: 2026-09-16 16:21:54 EDT → 2026-09-16 16:41:15 EDT · host jdmbair13m5 · claude (Opus 5)

What was broken

  1. jdmbair13m5 was off the tailnet. The Tailscale app (pid 969) and the network extension (pid 1000) were both running and the login-item helper was registered, so every process and launchctl check looked healthy. tailscale status said "Tailscale is stopped." — BackendState=Stopped, WantRunning=false, LoggedOut=false, node key held. Logged in, simply not connected. A stored preference, so it survives reboots.
  2. SSH from this host to rdmbair13m5/rdmbair15m5 was dead. A 2026-09-05 block in ~/.ssh/config pinned them to rdmbair13m5-1 / rdmbair15m5-1, duplicate tailnet nodes that no longer exist. Both failed with "Could not resolve hostname".
  3. /etc/hosts carried a fleet-lan-hosts block on all six Macs pinning bare hostnames to LAN IPs. It had rotted twice already and was wrong again: it claimed jdmbair13m5 was on 192.168.0.x while the host was on 192.168.1.225.
  4. SSH configs had drifted apart — rdmbair13m5 pinned rdmpw3275m to a cross-subnet LAN IP behind a ProxyJump; rdmpw3265m knew only rdmsm4x; three hosts had hand-rolled IP blocks.

What was done

Change Where
tailscale up --accept-routes — backend Stopped → Running, node online jdmbair13m5
Removed the fleet-lan-hosts block from /etc/hosts (backup /etc/hosts.bak-*) all 6
Replaced every hand-rolled ssh pin with one generated block: LAN first, Tailscale fallback all 6
Installed ~/bin/fleet-lan-probe (1-second TCP probe against the live LAN name) all 6
Installed com.eastcoastscience.tailscale-guard LaunchAgent (5 min + at load + on net change) all 6
Forced a fleet_ddns.sh run --force publish pass all 6
New fleet/net-naming/ with both scripts + README rdmsm4x (canonical)

ExitNodeAllowLANAccess went true → false on jdmbair13m5. It is inert without an exit node selected; tailscale up refuses the combination outright, which is what blocked the reconnect.

The naming scheme that replaces /etc/hosts

<host>.dataroo.net = live LAN IPv4, republished by fleet_ddns.sh every 5 min and on every network change. <host>.ts.dataroo.net = Tailscale. Bare <host> = Tailscale via MagicDNS. None of them is a pinned IP. ssh prefers the LAN answer when a 1-second probe says the peer is reachable there, and falls back to Tailscale otherwise.

Verified, not assumed

That test caught two real bugs before deployment: tailscale up wraps its refusal across two lines so a single-line glob never matches it, and the command it suggests already carries --timeout=, which is a hard error when passed twice.

Still open