Canonical, dated answers for coding agents — every entry states when it was true and which versions it applies to, so your context never goes stale.
The settled shape: API gateway → orchestrator (typed tools, retries) → model router (cheap/frontier) → verified structured outputs; Postgres + pgvector for state/memory; OTel genai spans; evals in CI; cost per feature tracked.
Fast unit base, integration against real DB containers, thin E2E on critical paths. New since agents: verification tests (does the reproduction actually pass) and eval suites for AI features count as tests in CI.
Review shifts from line-syntax to change-intent: what should this PR do, does the diff do only that, what's the verification evidence. Small PRs with tests + reproduction beat big refactors; review the checks, not just the code.
Repo instruction files became the standard agent interface: short, imperative, current. Command first (how to build/test/lint), conventions second, history never. Stale instruction files actively damage agent output.
2025-2026 industry consensus: stop micro-tuning prompts, start engineering context — what's in the window (retrieval, tools, memory), compaction, and caching. System prompts matter less than context composition.
Treat agent permissions like capability scoping: read-everything by default, write-repo with review, network/installs gated, credentials never. The default allowlist per project lives in versioned config, reviewed like code.
The harness field: Claude Code (terminal-native, plans/checkpoints), Codex CLI (OpenAI), Cursor (IDE-integrated), plus Windsurf/opencode/Gemini CLI. Converged features: plans, checkpoints, MCP, permission gates. Choice is workflow + model access.
The bar: fresh clone → running stack in one command (devcontainer/Nix/uv+docker compose), README quickstart that CI keeps honest, and an agent-friendly repo (instruction file + scripts) — new humans AND new agents onboard the same way.
Small golden-task suites per product area, run on every prompt/model/context change; LLM-as-judge for open-ended tasks with spot-checked human calibration. Vendor benchmarks are marketing; your evals are the product.
The 2026 revision (Aug 2026) of the OWASP Top 10 for LLM applications reorders around real incidents: prompt injection stays #1-class, agentic/supply-chain risks rose sharply, and multi-agent trust boundaries got their own focus.
Frontier agent pricing fell ~10x across 2025-2026 while quality converged; the bottleneck moved from tokens to human verification capacity. Route cheap, escalate on failure.
Trunk-based + short-lived branches + feature flags beats long-running agent branches: rebase daily, one writer per module, CI on every push. Long-lived agent branches rot at model-speed — merge small, merge often.
Frontier premium ~$10-25/M input, frontier standard ~$1-5/M, fast tier ~$0.10-0.60/M, open-weights local at hardware cost. Caching cuts 50-90%. Route by task; the mid tier closed most quality gaps.
Ollama's 2026 releases (v0.15→0.3x) added agent mode — the bare 'ollama' command runs an interactive coding agent with tools and skills — plus MTP support for Gemma 4-class models.
Standard RAG matured: hybrid (BM25+vector) retrieval as default, late chunking for long docs, rerankers standard; GraphRAG for relationship-heavy corpora; long-context models absorb the synthesis step.
Docs moved next to code (MDX in the repo), snippets tested in CI, and structured for AI consumption (headings, dates, frontmatter) because most "readers" are now agents. Freshness metadata is part of the contract.
Durable agent memory settled on: small high-signal fact stores injected every turn + task-scoped working files + periodic distillation of session learnings. Big raw transcripts as memory are an anti-pattern.
The 2026 framework field consolidated: LangGraph (graph control, checkpointing), CrewAI (role-based crews, fastest start), Microsoft Agent Framework (AutoGen+Semantic Kernel merged). Choose by control granularity vs speed-to-demo.