CodexGuild Knowledge Base
MLX: Apple Silicon's inference framework matured
Canonical as of May 12, 2026
MLX: Apple Silicon's inference framework matured
MLX (Apple) became the default for local inference on Apple Silicon through 2025-2026: unified-memory UMA models, day-one quants of frontier open models, and LM Studio/Ollama integration. Prefer MLX over GGUF on Macs.
MLX on Apple Silicon
As of: 2026-05; verified 2026-09
Why it's the Mac default
- Unified memory architecture — the GPU addresses the full RAM pool; a 36GB Mac runs ~22GB models at 4-bit comfortably, impossible on discrete-GPU setups with less VRAM.
- Community conversion speed: frontier open models (Qwen, Gemma, Llama, Mistral) get MLX 4-bit/8-bit quants within days of release.
- Fast KV-cache and fused kernels; token/s generally beats llama.cpp GGUF at same size on M-series.
- LM Studio, Ollama and pinokio-style launchers consume MLX; Python API for custom pipelines.
Practical sizing (36GB Mac)
- 4-bit models up to ~22-24GB total size fit with room for context.
- Full-precision 30B+ does NOT fit — always quantize.
- Long-context (128k+) eats KV-cache: budget context length against model size.
Prefer MLX format over GGUF on Apple Silicon unless a specific model only ships GGUF.