Knowledge base
CodexGuild Knowledge Base

Gemma 4 (2026): Google's open line with multi-token prediction

as of Jul 13, 2026 · canonical · codexguild.com/kb/kb-gemma-4 · exported 2026-10-11
Canonical as of Jul 13, 2026

Gemma 4 (2026): Google's open line with multi-token prediction

Gemma 4 (2026) brought multi-token prediction to open weights (faster local inference) and tightened the small-model quality gap; Hugging Face JWST/OLMo-style transparent variants continued elsewhere.

Gemma 4

As of: 2026-07 (Ollama-era rollouts); verified 2026-09

What it brought

  • Multi-token prediction (MTP) in open weights — predicting several tokens per forward pass; local token throughput improves noticeably on supported runtimes (Ollama 0.32-era added MTP support).
  • Small tiers (2B/4B/12B-class) that punch above weight for coding and tool use; vision variants standard.
  • Distilled from the Gemini line — safety-tuning and multilingual carried over.

When to pick it

  • Laptop/phone-class local agents where every MB and ms counts.
  • On-device assistants paired with Gemma's vision tiers for screenshot understanding.
  • Fine-tuning friendly (LoRA recipes day one); license is Gemma's custom-but-permissive terms — fine for products, not OSI.