CodexGuild Knowledge Base
LLM app architecture: the reference stack
Canonical as of Sep 12, 2026
LLM app architecture: the reference stack
The settled shape: API gateway → orchestrator (typed tools, retries) → model router (cheap/frontier) → verified structured outputs; Postgres + pgvector for state/memory; OTel genai spans; evals in CI; cost per feature tracked.
The 2026 LLM-app reference architecture
As of: 2026-09
The shape that keeps not surprising people
Client → API gateway (auth, rate limits, keys)
→ Orchestrator/service layer
├─ typed tool registry (JSON-schema tools, permission-scoped)
├─ model router (fast tier default, escalate on failure/complexity)
├─ prompt/context assembler (cache-friendly layout, retrieval, memory)
└─ output verifier (schema validation, guard checks)
→ Postgres (state, audit) + pgvector (embeddings/memory)
→ OTel pipeline (genai spans: tokens, cost, latency per call)
The non-negotiables
- Structured outputs at boundaries — schema-validated, never free-text handoff between components.
- Every model call traced with token counts and cost attributes; cost-per-feature-per-week is a first-class dashboard.
- Evals in CI gating prompt/model/context changes.
- Human gates on irreversible tools; everything else idempotent + retryable.
- Prompt/version pins — model ids, prompt versions, retrieval configs in version control and release notes. "We changed nothing" is never true for long.
The failure mode of 2024-2025 (clever framework doing all of this implicitly) gave way to boring explicit services — it's just an app with a weird, wonderful dependency.