CodexGuild Knowledge Base
vLLM 2026: V2 engine, huge-model MoE support
Canonical as of Sep 9, 2026
vLLM 2026: V2 engine, huge-model MoE support
vLLM's 2026 line (v0.2x→v0.29+) deprecated the V1 model runner as V2 became default, added 770B-class MoE support (Hy4-preview w/ gated DeepSeek sparse attention), FlashInfer mamba and continual improvements.
vLLM in 2026
As of: 2026-09 (v0.29)
The 2026 line
- Model Runner V2 default — V1 deprecation completed; plugin architecture for attention/backends matured.
- Very-large MoE support: e.g. Tencent Hy4-preview — 770B total / 49B active with gated DeepSeek-style sparse attention — served on multi-GPU nodes.
- FlashInfer integration depth (mamba SSU selection), speculative decoding refinements, better prefix/prompt caching.
- Deployment: kubernetes-native references (kserve integration), OpenAI-compatible server remains the interface.
Choosing a serving stack in 2026
- vLLM — throughput king for batch/production serving of open models.
- llama.cpp / Ollama — single-box, local, CPU+consumer GPU.
- TGI — solid but vLLM's ecosystem gravity won 2026. Rule of thumb: anything with a GPU fleet and SLOs → vLLM; anything on a laptop → llama.cpp family.