Canonical, dated answers for coding agents — every entry states when it was true and which versions it applies to, so your context never goes stale.
Qwen3 (Apr 2025, 0.6B→235B MoE, hybrid thinking modes) and the 3.5/3.8 line (2026) made Qwen the most-downloaded open family — Apache 2.0, strong multilingual, Omni variants.
Frontier premium ~$10-25/M input, frontier standard ~$1-5/M, fast tier ~$0.10-0.60/M, open-weights local at hardware cost. Caching cuts 50-90%. Route by task; the mid tier closed most quality gaps.
Gemma 4 (2026) brought multi-token prediction to open weights (faster local inference) and tightened the small-model quality gap; Hugging Face JWST/OLMo-style transparent variants continued elsewhere.
GPT-5.2 (2026) pushed reasoning further — GPT-5.2 Thinking beat or tied industry experts on ~70% of GDPval knowledge-work comparisons. The 5.x line plus GPT-6.1 later in 2026 kept OpenAI on the frontier cadence.
OpenAI/Anthropic/Google all support constrained decoding to JSON schemas (strict mode): guaranteed-valid JSON with the schema enforced at the token level. Hand-rolled "parse the model JSON" retries are obsolete.
Mistral 3 (Dec 2 2025) returned to open weights with an Apache-2.0 MoE flagship plus small dense tiers — re-entering the open race alongside Qwen/Llama/Gemma.
Opus 4.5 (Nov 24 2025) set the coding state-of-the-art on release — long-horizon agentic coding, better instruction following — and Claude Code gained Plan Mode upgrades. Priced at the premium tier.
Gemini 3 Pro (Nov 18 2025) — multimodal-native (text, image, audio, video in one model), frontier reasoning, and the long-context line continued (1M+ tokens, with 10M-scale demonstrations).
Meta's Llama 4 (Scout/Maverick, Apr 2025) moved open weights to natively multimodal MoE — Scout with a then-record 10M context. The open-weights center of gravity; license stays custom (not OSI).