rdmsm4x-changelog-20260830-0030-apple-mlx-suite-and-mcp-integration
rdmsm4x-changelog-20260830-0030-apple-mlx-suite-and-mcp-integration
Summary
Created canonical apple-mlx project at
~/dev/apps/apple-mlx, built automated Apple Silicon
installer and fleet deployment automation, implemented Model Context
Protocol (MCP) server mlx-mcp with JSON-RPC 2.0 stdio,
configured LaunchAgent daemon, and published comprehensive LLM runtime
ranking (Ollama vs MLX vs LM Studio vs Apple Foundation Models).
Details of Changes
1. New Canonical
Project: ~/dev/apps/apple-mlx
- Location:
rdmsm4x:~/dev/apps/apple-mlx/(Git repository initialized onmain). - Core Binaries:
bin/mlx-mcp&bin/mlx_mcp_server.py: Native stdio Model Context Protocol (MCP) server exposingmlx_generate,mlx_chat,mlx_list_models,mlx_load_model, andmlx_metal_status.bin/mlx-server: OpenAI-compatible HTTP server (/v1/chat/completions) wrappingmlx_lm.serverwith Metal GPU acceleration on port 8080.bin/mlx-cli: Fast interactive/non-interactive CLI text generation.- Symlinked to
~/bin/mlx-mcp,~/bin/mlx-server, and~/bin/mlx-cli.
2. Fleet Installation & Deployment Automation
scripts/install_mlx.zsh: Dynamic Homebrew detection (brew --prefix),uvvirtual environment creation (~/.local/share/mlx-env), package installation (mlx==0.32.2,mlx-lm==0.31.3,mlx-whisper,fastapi,uvicorn), Metal verification, and MCP schema registration.scripts/fleet_deploy_mlx.zsh: Multi-host deployment orchestrator across all 6 fleet machines (rdmsm4x,rdmbair15m5,rdmbair13m5,jdmbair13m5,rdmpw3265m,rdmpw3275m) over Tailscale SSH.launchd/com.eastcoastscience.mlx.plist: Native macOS LaunchAgent for persistent background model serving.
3. Comprehensive LLM Runtime Comparison
- Published
docs/LLM_RUNTIME_COMPARISON.md:- Rank 1: Apple MLX (
mlx-lm) — optimal memory bandwidth (94-98%), zero-copy unified memory, fastest TTFT and highest decode tokens/sec. - Rank 2: Ollama — best for multi-node fleet CLI ergonomics and rapid model pulling.
- Rank 3: omlx — MLX Metal speed with Ollama REST API compatibility.
- Rank 4: llama.cpp — bare-metal C++ GGUF inference.
- Rank 5: LM Studio — interactive desktop GUI testing.
- Apple Official Releases: Documented Apple ML
Research releases (MLX, OpenELM, Ferret-v2, Depth Pro, MobileCLIP,
mlx-community) vs on-device Apple Intelligence Foundation Models (AFM-on-Device 3B CoreML model).
- Rank 1: Apple MLX (
4. Verification & Testing
- Unit test suite
tests/test_mlx_mcp.pyexecuted: 4/4 tests passed (0.16s). - Verified Metal GPU compute on
rdmsm4x:Device(gpu, 0)initialized, zero-copy arrays evaluated. - Verified MCP stdio handshake, tools listing, and
mlx_metal_statustool calls. - Resolved fleet ticket
FEAT-20260830-40.