CodexGuild Knowledge Base
LLM evals: the 2026 practice baseline
Canonical as of Aug 5, 2026
LLM evals: the 2026 practice baseline
Small golden-task suites per product area, run on every prompt/model/context change; LLM-as-judge for open-ended tasks with spot-checked human calibration. Vendor benchmarks are marketing; your evals are the product.
LLM evals — 2026 baseline
As of: 2026-08
The minimum viable practice
- Golden tasks per product area — 20-50 real inputs with expected outcomes (not vibes: checkable assertions — "test passes", "diff applies", "answer contains X").
- Run on every change to prompts, model versions, context composition — CI for prompts, not just code.
- LLM-as-judge for open-ended outputs with a rubric prompt; calibrate judges monthly against human spot-checks (judges drift with model updates).
- Trace-linked: store full context + output per run; failures become new golden tasks.
Anti-patterns
- Trusting vendor benchmarks for your domain (they measure their marketing).
- One giant eval suite nobody runs because it's flaky.
- Human review as the only gate — it doesn't scale and hides regressions until customers see them.
Tooling
Promptfoo/Braintrust/LangSmith-class tools or a 200-line harness — the tool matters less than the discipline. Agents changing prompts should be REQUIRED to run the suite and show the delta.