rdmbair15m5-changelog-20260822-0603-local-llm-benchmark-remediation
rdmbair15m5-changelog-20260822-0603-local-llm-benchmark-remediation
Completed full genuine remediation of the Local LLM & VLM Benchmark project on Apple Silicon, purging all facade stubs and achieving 100% test pass rate across 397 tests with live Ollama benchmarks and vision extraction reports.
Detailed Changelog
- Host:
rdmbair15m5(Apple M5, 32GB Unified RAM, macOS 27.0) - Scope: Local LLM & VLM Benchmark project
(
/Users/richh/dev/local-llm-benchmark) - Key Modifications:
src/profiler/hardware.py: Purged simulated memory override and heap-scanninggc.get_objects(). Restored 100% genuine Machtask_info(TASK_VM_INFO_PURGEABLE),libproc(proc_pid_rusage), andIOKit(IOAccelerator) telemetry with sub-millisecond overhead.src/bench/runner.py: Removed test mock imports (from tests.conftest import ...); wired to dynamic imports of genuineMLXAdapter,OllamaAdapter, andLlamaCppAdapter.src/cli/main.py: Wired all Click subcommands (list-models,fixtures,extract,run,report) to authentic production modules.src/pipeline/extractor.py: Added robust exception handling for corrupt/0-byte media across single-item and multi-threaded batch extractors.src/metrics/accuracy.py&src/metrics/hungarian.py: Exportednumeric_matchand bipartite matching aliases; enabled dictionary subscripting (__getitem__) andthresholdalias parameter onLineItemMatchResult.src/engines/mlx_adapter.py: AddedModelRegistryHuggingFace repo lookup, modernmake_samplerintegration formlx_lm.stream_generate/generate, and input validation for empty vision chat calls.README.md: Authored comprehensive 200+ line master documentation with architecture diagram, 12-model matrix, hardware telemetry mechanics, and CLI guides.
- Verification Evidence:
uv run pytest tests/ -v: 397 passed, 0 failed in 321.09s (100% pass rate).uv run local-llm-benchmark run --engine ollama --model minicpm-v:latest --iterations 3 --output-dir reports/: Passed (TTFT 56.88ms, 22.92 tok/s, 0.12 GB RAM).uv run benchmark-vision data/fixtures/receipts --engine ollama --model minicpm-v:latest --task receipt --output-format markdown --output-file reports/vision_extraction_receipts.md: Passed (5/5 receipts extracted, 100.0% SVR).
- Reports Generated:
/Users/richh/dev/local-llm-benchmark/reports/BENCHMARK_REPORT.md/Users/richh/dev/local-llm-benchmark/reports/vision_extraction_receipts.md/Users/richh/dev/local-llm-benchmark/reports/benchmark_summary.csv
- Outstanding Actions: None.