rdmbair15m5-changelog-20260822-0640-local-llm-benchmark-suite
rdmbair15m5-changelog-20260822-0640-local-llm-benchmark-suite
Benchmarking suite and multimodal vision extraction framework for Apple Silicon (MLX, llama.cpp Metal, Ollama).
Scope
- Host: rdmbair15m5
- Path: ~/dev/local-llm-benchmark
Summary of Changes
- Implemented unified engine adapters for Ollama, Apple MLX (mlx-lm/mlx-vlm), and native llama.cpp Metal.
- Provisioned 12-model matrix (4 VLMs, 8 LLMs) optimized for a 32GB unified memory budget.
- Created synthetic and curated multimodal test datasets covering receipts/invoices, multi-page PDFs, and visual charts.
- Implemented accuracy metrics engine (CER, WER, NCED, Hungarian bipartite line-item matching with cent tolerance).
- Developed batch vision extraction pipeline with 4-tier JSON repair and CLI commands ('local-llm-benchmark', 'benchmark-vision', 'local-vlm-extract').
- Built Apple Silicon telemetry profiler (CTypes Mach task_info, libproc, IOKit, NSProcessInfo).
- Added multi-format benchmark reporting (Markdown, CSV, HTML interactive dashboard, Matplotlib PNG charts).
- Built 5-tier test suite with 402 passing tests verified via independent victory audit.
Verification Evidence
- 'uv run pytest -v': 402 passed, 0 failures.
- Victory Audit Verdict: VICTORY CONFIRMED.