oncall-irm
Route alerts, run on-call rotations, and drive incidents in Grafana IRM / OnCall — integrations (Alertmanager / Grafana Alerting / generic webhook / PagerDuty), Jinja2 routing + grouping templates, escalation chains (wait → notify schedule → notify team → webhook → auto-resolve), schedules (web + iC
- 0
- Installs
- —
- Rating
- —
- Success rate
- 3
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 2c1d741cd1aee408… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Grafana OnCall & IRM
OnCall docs: https://grafana.com/docs/oncall/latest/ IRM docs: https://grafana.com/docs/grafana-cloud/alerting-and-irm/
Grafana OnCall OSS is in maintenance mode (archived March 2026). Cloud users → IRM. Concepts (chains, schedules, integrations) are identical.
Prerequisites
- Grafana Cloud stack with IRM/OnCall enabled
- API token (
Authorization: <token>) - Slack workspace + admin to install the OnCall app (for ChatOps)
Core concepts
| Concept | Description |
|---|---|
| Integration | Webhook URL accepting alerts; one per source |
| Route | Jinja2 condition that maps to an escalation chain (first True wins) |
| Escalation Chain | Wait / notify schedule / notify team / webhook / auto-resolve steps |
| Schedule | Calendar-based rotation (web / iCal / Terraform) |
| Alert Group | Related alerts collapsed by a Grouping ID template |
| Notification Policy | Per-user channels — Slack, mobile push, SMS, phone, email |
Flow: alert → integration → routing template → escalation chain → notifications → ack / resolve.
Common Workflows
1. Wire Alertmanager → IRM and verify routing
# 1. In IRM: New integration → Alertmanager. Copy the webhook URL.
# 2. alertmanager.yml (see references/integrations.md for full block):
receivers:
- name: grafana-oncall
webhook_configs:
- url: https://<stack>.grafana.net/integrations/v1/alertmanager/<id>/
send_resolved: true
max_alerts: 100
# 3. Test the routing template BEFORE going live (UI: Integration → Route → "Preview")
# Paste a sample payload (use a real Alertmanager test webhook). Expect:
# Routing result: True
# Escalation chain selected: <expected chain>
# If False — your Jinja expression is wrong; fix and re-preview.
# 4. End-to-end test — fire amtool (or any test webhook) at the URL
amtool alert add foo severity=critical team=platform --alertmanager.url http://localhost:9093
# 5. Verify in IRM → Alert Groups (should appear within ~5s) — confirm:
# - Correct route fired
# - Correct schedule/user notified
# - Slack message appeared in the configured channel
Routing template syntax + Jinja helpers: references/templates-schedules.md.
Other integrations (Grafana Alerting, generic webhook, Slack): references/integrations.md.
2. Build an escalation chain + verify
1. Notify "Primary On-Call" (Important Notifications)
2. Wait 5 min
3. Notify "Primary On-Call" (Default Notifications)
4. Wait 10 min
5. Notify team "Platform"
6. Trigger outgoing webhook (PagerDuty / ticket)
# Verify: create a test alert (UI → Integration → "Send demo alert"),
# then watch the alert-group timeline tick through steps 1 → 6.
# Use IRM → Escalation Chains → "Test" if available, otherwise the demo alert is the canonical check.
3. Create a schedule from iCal + verify "who is on-call right now"
# 1. Create the schedule
curl -X POST https://<stack>.grafana.net/api/v1/schedules/ \
-H "Authorization: <token>" -H "Content-Type: application/json" \
-d '{"name":"Platform On-Call",
"ical_url_primary":"https://calendar.example.com/platform.ics",
"slack":{"channel_id":"C123456ABC","user_group_id":"S123456ABC"}}'
# 2. Verify the schedule was created
curl -s https://<stack>.grafana.net/api/v1/schedules/ \
-H "Authorization: <token>" | jq '.results[] | select(.name=="Platform On-Call")'
# 3. Verify who is on-call right now
curl -s https://<stack>.grafana.net/api/v1/schedules/<schedule_id>/next_shifts/ \
-H "Authorization: <token>" | jq '.results[0]'
# Expect a shift starting now (or recently) with the right user_id.
Terraform variant + shift block: references/templates-schedules.md.
Best practices
- Keep chains ≤4 levels with a definitive final step (webhook to PagerDuty or auto-resolve)
- Always set
send_resolved: truein Alertmanager so OnCall auto-resolves - Use
max_alerts: 100in Alertmanager webhook config - Combine Slack + mobile push for delivery reliability
- Assign integrations/schedules to teams for RBAC
Resources
Files
3- SKILL.md
4e115f1f7f5.4 KB - references/integrations.md
146ac272852.1 KB - references/templates-schedules.md
91516b38533.2 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from grafana/skills8
Cut Grafana Cloud Metrics cost by shrinking active-series count with Adaptive Metrics aggregation rules — auto-recommendations from query history, custom exact/regex rules, label-drop config, unused-metric detection, and Alloy remote_write fallback. Use when investigating a high Mimir/Grafana Cloud
Manage Grafana Cloud accounts — organizations, stacks, RBAC roles and assignments, SSO/SAML/OAuth/GitHub auth, service accounts for CI/CD, user invites, team membership, and API-driven provisioning. Creates stacks via the Cloud API, mints service-account tokens, applies role assignments, configures
Use when the user asks to "write a validator", "add validation", "implement admission control", "write a mutating webhook", "add a mutation handler", "validate incoming resources", "implement admission logic", "add admission webhooks", "write ingress validation", or asks how to validate or mutate re
Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook), notification policies with hierarchical matchers, silences, mute timings, on-call schedules and escala
Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo / Pyroscope. Covers the Alloy config language (blocks, `sys.env`, component refs), `prometheus.scrape`
Get RED metrics + service maps + frontend RUM + AI/LLM monitoring out of Grafana Cloud — Application Observability (`traces_spanmetrics_*` from OTel traces, p50/p95/p99 latency, exemplar-to-trace, traces-to-logs / profiles), Frontend Observability with the Faro Web SDK (Core Web Vitals, session repl
Use when starting any grafana-app-sdk work — scaffolding a Grafana app, initializing a Grafana App Platform app, picking a deployment mode (standalone operator / grafana/apps / frontend-only), wiring app-specific config, or onboarding to the SDK. Covers `grafana-app-sdk` CLI install, `project init`
Connect AI coding agents (Claude Code, Cursor, VS Code, OpenAI Codex) to Grafana Cloud via the `mcp-grafana` Model Context Protocol server. Installs the server with `go install`, generates a Grafana service-account token, wires `~/.claude/settings.json` or `~/.cursor/mcp.json` with the `command` + `
Related backend skillsscan passed
PostHog integration for FastAPI applications
Verify a local agent API, temporary gateway tunnel, and remote sandbox callback with a tool-free task, then restore the original app connection.
Report browser/API/CLI/job/worker/webhook bugs. (gstack)
This skill should be used when the user wants to "package an MCP server", "bundle an MCP", "make an MCPB", "ship a local MCP server", "distribute a local MCP", discusses ".mcpb files", mentions bundling a Node or Python runtime with their MCP server, or needs an MCP server that interacts with the lo
Identifies external providers, merchants, nonprofits, platforms, APIs, and software services, and resolves the documented way to engage them — to pay, donate, subscribe, book, provision, or integrate with them. MUST be used BEFORE web search, model memory, or any other directory/vendor-lookup skill
Mount tRPC as a Fastify plugin with fastifyTRPCPlugin from @trpc/server/adapters/fastify. Configure prefix, trpcOptions (router, createContext, onError). Enable WebSocket subscriptions with useWSS and @fastify/websocket. Set routerOptions.maxParamLength for batch requests. Requires Fastify v5+. Fast