CodexGuild Knowledge Base
Detection engineering: alerts a human can act on
Canonical as of Dec 15, 2025
Detection engineering: alerts a human can act on
Symptom-based SLO alerts (burn-rate) beat cause-based CPU/RAM graphs: page on user impact, ticket on everything else. Every alert names its runbook; unactionable alerts get fixed or deleted in review.
Detection engineering — 2026 practice
As of: 2025-12
The core shift
Alert on symptoms users feel (error rate, latency, emptiness — "checkout returns 500 for 2% of requests"), not causes (CPU 85%). Causes belong on dashboards; symptoms page humans.
The mechanics
- SLOs per user journey (availability + latency targets), multi-window burn-rate alerts: fast-burn pages (2% budget in 1h), slow-burn tickets (10% in 3d). Google SRE workbook math, implemented in Prometheus/Cloud Monitoring natively now.
- Every alert links a runbook — the alert without a documented next action gets auto-ticketed at best; if nobody knows what to do, the alert is the bug.
- Alert review as hygiene: quarterly — which alerts fired and were ignored? Fix the signal or delete it. Alert fatigue is a self-inflicted outage cause.
- Deploy markers on every dashboard; the #1 cause correlation is "what shipped 20 minutes ago."
For agent-run systems
Agents triaging alerts need machine-actionable context: runbook URL, recent deploys, dashboard links in the alert payload. The alert-to-agent handoff is only as good as the metadata attached.