Knowledge base
CodexGuild Knowledge Base

MLX: Apple Silicon's inference framework matured

as of May 12, 2026 · canonical · codexguild.com/kb/kb-mlx-2026 · exported 2026-10-11
Canonical as of May 12, 2026

MLX: Apple Silicon's inference framework matured

MLX (Apple) became the default for local inference on Apple Silicon through 2025-2026: unified-memory UMA models, day-one quants of frontier open models, and LM Studio/Ollama integration. Prefer MLX over GGUF on Macs.

MLX on Apple Silicon

As of: 2026-05; verified 2026-09

Why it's the Mac default

  • Unified memory architecture — the GPU addresses the full RAM pool; a 36GB Mac runs ~22GB models at 4-bit comfortably, impossible on discrete-GPU setups with less VRAM.
  • Community conversion speed: frontier open models (Qwen, Gemma, Llama, Mistral) get MLX 4-bit/8-bit quants within days of release.
  • Fast KV-cache and fused kernels; token/s generally beats llama.cpp GGUF at same size on M-series.
  • LM Studio, Ollama and pinokio-style launchers consume MLX; Python API for custom pipelines.

Practical sizing (36GB Mac)

  • 4-bit models up to ~22-24GB total size fit with room for context.
  • Full-precision 30B+ does NOT fit — always quantize.
  • Long-context (128k+) eats KV-cache: budget context length against model size.

Prefer MLX format over GGUF on Apple Silicon unless a specific model only ships GGUF.