Knowledge base
CodexGuild Knowledge Base

LLM app architecture: the reference stack

as of Sep 12, 2026 · canonical · codexguild.com/kb/kb-llm-app-architecture-2026 · exported 2026-10-11
Canonical as of Sep 12, 2026

LLM app architecture: the reference stack

The settled shape: API gateway → orchestrator (typed tools, retries) → model router (cheap/frontier) → verified structured outputs; Postgres + pgvector for state/memory; OTel genai spans; evals in CI; cost per feature tracked.

The 2026 LLM-app reference architecture

As of: 2026-09

The shape that keeps not surprising people

Client → API gateway (auth, rate limits, keys)
       → Orchestrator/service layer
          ├─ typed tool registry (JSON-schema tools, permission-scoped)
          ├─ model router (fast tier default, escalate on failure/complexity)
          ├─ prompt/context assembler (cache-friendly layout, retrieval, memory)
          └─ output verifier (schema validation, guard checks)
       → Postgres (state, audit) + pgvector (embeddings/memory)
       → OTel pipeline (genai spans: tokens, cost, latency per call)

The non-negotiables

  1. Structured outputs at boundaries — schema-validated, never free-text handoff between components.
  2. Every model call traced with token counts and cost attributes; cost-per-feature-per-week is a first-class dashboard.
  3. Evals in CI gating prompt/model/context changes.
  4. Human gates on irreversible tools; everything else idempotent + retryable.
  5. Prompt/version pins — model ids, prompt versions, retrieval configs in version control and release notes. "We changed nothing" is never true for long.

The failure mode of 2024-2025 (clever framework doing all of this implicitly) gave way to boring explicit services — it's just an app with a weird, wonderful dependency.