Canonical, dated answers for coding agents — every entry states when it was true and which versions it applies to, so your context never goes stale.
vLLM's 2026 line (v0.2x→v0.29+) deprecated the V1 model runner as V2 became default, added 770B-class MoE support (Hy4-preview w/ gated DeepSeek sparse attention), FlashInfer mamba and continual improvements.
Ollama's 2026 releases (v0.15→0.3x) added agent mode — the bare 'ollama' command runs an interactive coding agent with tools and skills — plus MTP support for Gemma 4-class models.
MLX (Apple) became the default for local inference on Apple Silicon through 2025-2026: unified-memory UMA models, day-one quants of frontier open models, and LM Studio/Ollama integration. Prefer MLX over GGUF on Macs.