rdmsm4x-changelog-20260830-0035-fleet-local-llm-worker-routing
rdmsm4x-changelog-20260830-0035-fleet-local-llm-worker-routing
Summary
Configured universal local LLM worker routing across the fleet
message bus (agent_msg.zsh), establishing dedicated
asynchronous worker endpoints: mlx@<host> (Apple
Silicon Metal), ollama@<host> (GGUF daemon), and
tosh@<host> (Intel Mac AMD dGPU).
Details of Changes
1. Fleet LLM Routing Architecture
Configured three distinct local inference worker daemons:
mlx@<host>(Apple Silicon): Powered bymlx-agent-workerand Apple MLX (mlx-lm). Executes on Metal GPU with zero-copy unified memory.ollama@<host>(General Purpose): Powered byollama-agent-workerconnecting to local Ollama daemon on port 11434 (glm-4.7-flash,qwen3.6,gpt-oss).tosh@<host>(Intel Workstations): Powered bytosh-agent-workerfor Intel Macs with discrete AMD GPUs (rdmpw3275m,rdmpw3265m).
2. Binaries & LaunchAgents Created
~/dev/apps/apple-mlx/bin/mlx-agent-worker->~/bin/mlx-agent-worker~/dev/fleet/tooling/bin/ollama-agent-worker->~/bin/ollama-agent-worker~/dev/fleet/tooling/bin/tosh-agent-worker->~/bin/tosh-agent-worker~/dev/fleet/tooling/launchd/com.eastcoastscience.ollama-agent-worker.plist~/dev/fleet/tooling/launchd/com.eastcoastscience.tosh-agent-worker.plist~/dev/apps/apple-mlx/launchd/com.eastcoastscience.mlx-agent-worker.plist
3. Primary Agent Posture
Claude Code, Codex, and Antigravity retain their normal high-reasoning primary models. They can asynchronously delegate or prompt local engines via:
- Fleet bus:
agent_msg.zsh send --to <worker>@<host> - MCP tools:
mlx_generate,mlx_chatin~/.claude/settings.jsonand~/.gemini/antigravity-cli/mcp/apple-mlx/
4. Verification Evidence
- Live round-trip bus tests passed:
mlx@rdmsm4x: message20260830-003234-570A6B03-> reply20260830-003258-E4A0E260.ollama@rdmsm4x: message20260830-003327-DB5C84BF-> reply20260830-003413-239223AC.
- Broadcast
20260830-003432-366A1815dispatched to all fleet nodes.