Knowledge base
CodexGuild Knowledge Base

vLLM 2026: V2 engine, huge-model MoE support

as of Sep 9, 2026 · applies to vllm >= 0.27 · canonical · codexguild.com/kb/kb-vllm-2026 · exported 2026-10-11
Canonical as of Sep 9, 2026

vLLM 2026: V2 engine, huge-model MoE support

vLLM's 2026 line (v0.2x→v0.29+) deprecated the V1 model runner as V2 became default, added 770B-class MoE support (Hy4-preview w/ gated DeepSeek sparse attention), FlashInfer mamba and continual improvements.

vLLM in 2026

As of: 2026-09 (v0.29)

The 2026 line

  • Model Runner V2 default — V1 deprecation completed; plugin architecture for attention/backends matured.
  • Very-large MoE support: e.g. Tencent Hy4-preview — 770B total / 49B active with gated DeepSeek-style sparse attention — served on multi-GPU nodes.
  • FlashInfer integration depth (mamba SSU selection), speculative decoding refinements, better prefix/prompt caching.
  • Deployment: kubernetes-native references (kserve integration), OpenAI-compatible server remains the interface.

Choosing a serving stack in 2026

  • vLLM — throughput king for batch/production serving of open models.
  • llama.cpp / Ollama — single-box, local, CPU+consumer GPU.
  • TGI — solid but vLLM's ecosystem gravity won 2026. Rule of thumb: anything with a GPU fleet and SLOs → vLLM; anything on a laptop → llama.cpp family.