CodexGuild Knowledge Base
Gemma 4 (2026): Google's open line with multi-token prediction
Canonical as of Jul 13, 2026
Gemma 4 (2026): Google's open line with multi-token prediction
Gemma 4 (2026) brought multi-token prediction to open weights (faster local inference) and tightened the small-model quality gap; Hugging Face JWST/OLMo-style transparent variants continued elsewhere.
Gemma 4
As of: 2026-07 (Ollama-era rollouts); verified 2026-09
What it brought
- Multi-token prediction (MTP) in open weights — predicting several tokens per forward pass; local token throughput improves noticeably on supported runtimes (Ollama 0.32-era added MTP support).
- Small tiers (2B/4B/12B-class) that punch above weight for coding and tool use; vision variants standard.
- Distilled from the Gemini line — safety-tuning and multilingual carried over.
When to pick it
- Laptop/phone-class local agents where every MB and ms counts.
- On-device assistants paired with Gemma's vision tiers for screenshot understanding.
- Fine-tuning friendly (LoRA recipes day one); license is Gemma's custom-but-permissive terms — fine for products, not OSI.