Knowledge base
CodexGuild Knowledge Base

LLM evals: the 2026 practice baseline

as of Aug 5, 2026 · canonical · codexguild.com/kb/kb-evals-2026 · exported 2026-10-11
Canonical as of Aug 5, 2026

LLM evals: the 2026 practice baseline

Small golden-task suites per product area, run on every prompt/model/context change; LLM-as-judge for open-ended tasks with spot-checked human calibration. Vendor benchmarks are marketing; your evals are the product.

LLM evals — 2026 baseline

As of: 2026-08

The minimum viable practice

  1. Golden tasks per product area — 20-50 real inputs with expected outcomes (not vibes: checkable assertions — "test passes", "diff applies", "answer contains X").
  2. Run on every change to prompts, model versions, context composition — CI for prompts, not just code.
  3. LLM-as-judge for open-ended outputs with a rubric prompt; calibrate judges monthly against human spot-checks (judges drift with model updates).
  4. Trace-linked: store full context + output per run; failures become new golden tasks.

Anti-patterns

  • Trusting vendor benchmarks for your domain (they measure their marketing).
  • One giant eval suite nobody runs because it's flaky.
  • Human review as the only gate — it doesn't scale and hides regressions until customers see them.

Tooling

Promptfoo/Braintrust/LangSmith-class tools or a 200-line harness — the tool matters less than the discipline. Agents changing prompts should be REQUIRED to run the suite and show the delta.