rdmbair15m5-changelog-20260822-0520-local-llm-benchmark-m2-model-matrix
rdmbair15m5-changelog-20260822-0520-local-llm-benchmark-m2-model-matrix
Implemented Milestone 2 for Local LLM & VLM Benchmark:
comprehensive 32GB Unified Memory Model Matrix & Registry
(src/models/registry.py), local availability and pre-flight
headroom provisioner (src/models/provisioner.py), module
exports (src/models/__init__.py), and full unit test suite
(tests/test_models.py).
Scope
- Host:
rdmbair15m5(Apple M5, 32GB Unified RAM, macOS 27.0) - Project:
~/dev/local-llm-benchmark - Milestone: Milestone 2 (Model Registry & Provisioning Matrix)
Files Modified / Created
src/models/registry.py: ModelType, ModalityType, QuantizationType, ModelFamily enums; ModelMetadata dataclass with 4-bit quantization and 75% macOS GPU ceiling validation; 12 canonical models (4 VLMs: Llama 3.2 Vision 11B, MiniCPM-V 2.6 8B, Qwen2.5-VL 7B, Moondream2; 8 LLMs: Qwen 2.5 14B/32B, DeepSeek-R1 14B/32B, Gemma 2 9B/27B, Qwen 2.5 Coder 14B/32B); case-insensitive lookup and filter functions.src/models/provisioner.py: ModelProvisioner class and functions for Ollama/api/tagsverification, HuggingFace Hub cache inspection (~/.cache/huggingface/hub), local GGUF search, host virtual memory preflight checks, and model download helpers.src/models/__init__.py: Clean public API export.tests/test_models.py: 21 comprehensive unit tests covering enums, metadata, 12 canonical model specs, fuzzy/exact lookups, memory budget boundary conditions, and mock-assisted provisioner functions.
Verification Evidence
uv run pytest tests/test_models.py: 21 passed in 0.11s.uv run pytest: 292 passed in 2.75s (full repository test suite passing with 0 regressions).