skills/ grafana/skills

ml-ai

Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL/LogQL/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN outlier detection), Sift (8-analysis automated root-cause), Knowledge Graph + RCA Workbench, and the LLM

0
Installs
—
Rating
—
Success rate
3
Files scanned
Scan passedai-ml
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

3 files scannedscanner v1.2.0Oct 10, 2026

Content sha256 29bb8502de6cda4e… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Grafana Cloud AI & ML

Docs: https://grafana.com/docs/grafana-cloud/alerting-and-irm/machine-learning/

ML alerting + automated RCA + LLM-powered Assistant in one Grafana Cloud stack.

Prerequisites

  • Grafana Cloud stack (Pro / Advanced — most features GA, some in preview)
  • API token with plugins:write for ML / Sift / LLM-plugin endpoints
  • For Dynamic Alerting: at least 14 days (ideally 90d) of history for the metric you want to forecast

Common Workflows

1. Forecasting alert with Dynamic Alerting

# 1. Create forecast job (Prophet — learns daily/weekly seasonality)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-ml-app/resources/ml/v1/forecast \
  -H "Authorization: Bearer <token>" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "cpu-forecast",
    "metric": "avg(rate(node_cpu_seconds_total{mode=\"user\"}[5m]))",
    "datasourceId": 1,
    "interval": 300,
    "trainingWindow": "90d",
    "forecastWindow": "7d",
    "algorithm": { "name": "prophet", "config": {} }
  }'

# 2. Verify job is producing the predicted-value metric (may take a few minutes).
#    <datasourceId> must match the datasourceId used above (find it via
#    GET /api/datasources), or run the query from Explore instead.
curl -s -H "Authorization: Bearer <token>" \
  'https://<stack>.grafana.net/api/datasources/proxy/<datasourceId>/api/v1/query?query=ml_forecast_upper{job="cpu-forecast"}' \
  | jq '.data.result | length'
# Expect > 0

# 3. Add an alert that fires when actual exceeds the upper bound
# expr:  avg(rate(node_cpu_seconds_total{mode="user"}[5m]))
#         > ml_forecast_upper{job="cpu-forecast"} * 1.1

2. Outlier alert — one service deviates from peers

# 1. Create outlier job (DBSCAN — groups peers, flags the odd one)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-ml-app/resources/ml/v1/outlier \
  -H "Authorization: Bearer <token>" -H "Content-Type: application/json" \
  -d '{
    "name": "service-error-outliers",
    "metric": "sum(rate(http_requests_total{status=~\"5..\"}[5m])) by (service)",
    "datasourceId": 1,
    "interval": 300,
    "algorithm": { "name": "dbscan", "sensitivity": 0.5, "config": { "epsilon": 0.5 } }
  }'

# 2. Verify the score metric exists (<datasourceId> must match the
#    datasourceId used above, or run the query from Explore instead)
curl -s -H "Authorization: Bearer <token>" \
  'https://<stack>.grafana.net/api/datasources/proxy/<datasourceId>/api/v1/query?query=ml_outlier_score{job="service-error-outliers"}' \
  | jq '.data.result | length'

# 3. Alert when ml_outlier_score{job="service-error-outliers"} > 0.8 for 5m

3. Run a Sift investigation

# 1. Trigger from API (or from Explore / Incident / OnCall)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-sift-app/resources/sift/v1/investigations \
  -H "Authorization: Bearer <token>" -H "Content-Type: application/json" \
  -d '{ "name":"checkout-spike","start":"2024-02-01T10:00:00Z","end":"2024-02-01T10:30:00Z",
        "filters":{"service":"checkout","namespace":"production"} }'

# 2. The response includes an investigation ID — open it in the UI:
#    https://<stack>.grafana.net/a/grafana-sift-app/investigations/<id>
# 3. Verify analyses ran — each of the 8 checks shows ✔ or ✖ with linked evidence.

See references/sift.md for the full 8-analysis table.

4. Wire up the LLM Plugin

# 1. Provision (provisioning/plugins/llm.yaml — see references/llm-and-graph.md)
apiVersion: 1
apps:
  - type: grafana-llm-app
    jsonData: { openAIUrl: https://api.openai.com, openAIModel: gpt-4o }
    secureJsonData: { openAIKey: sk-... }
# 2. Restart Grafana, then verify the health endpoint reports the configured provider
curl -s -H "Authorization: Bearer <token>" \
  https://<stack>.grafana.net/api/plugins/grafana-llm-app/health | jq
# Expect: {"status":"ok", ...}

# 3. Verify in a panel — open any panel, click the Assistant icon, ask "what does this query do?"

See references/llm-and-graph.md for Assistant capabilities, Knowledge Graph search syntax, and Adaptive Metrics recommendations.

Resources

Files

3
9.1 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from grafana/skills8

adaptive-metrics

Cut Grafana Cloud Metrics cost by shrinking active-series count with Adaptive Metrics aggregation rules — auto-recommendations from query history, custom exact/regex rules, label-drop config, unused-metric detection, and Alloy remote_write fallback. Use when investigating a high Mimir/Grafana Cloud

Scan passed 0
admin

Manage Grafana Cloud accounts — organizations, stacks, RBAC roles and assignments, SSO/SAML/OAuth/GitHub auth, service accounts for CI/CD, user invites, team membership, and API-driven provisioning. Creates stacks via the Cloud API, mints service-account tokens, applies role assignments, configures

Scan passed 0
admission-control

Use when the user asks to "write a validator", "add validation", "implement admission control", "write a mutating webhook", "add a mutation handler", "validate incoming resources", "implement admission logic", "add admission webhooks", "write ingress validation", or asks how to validate or mutate re

Scan passed 0
alerting-irm

Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack/PagerDuty/email/webhook), notification policies with hierarchical matchers, silences, mute timings, on-call schedules and escala

Scan passed 0
alloy

Build a unified telemetry pipeline with Grafana Alloy — one OpenTelemetry-compatible binary that collects metrics, logs, traces, and profiles and ships to Grafana Cloud / Prometheus / Loki / Tempo / Pyroscope. Covers the Alloy config language (blocks, `sys.env`, component refs), `prometheus.scrape`

Scan passed 0
app-observability

Get RED metrics + service maps + frontend RUM + AI/LLM monitoring out of Grafana Cloud — Application Observability (`traces_spanmetrics_*` from OTel traces, p50/p95/p99 latency, exemplar-to-trace, traces-to-logs / profiles), Frontend Observability with the Faro Web SDK (Core Web Vitals, session repl

Needs review 0
app-sdk-concepts

Use when starting any grafana-app-sdk work — scaffolding a Grafana app, initializing a Grafana App Platform app, picking a deployment mode (standalone operator / grafana/apps / frontend-only), wiring app-specific config, or onboarding to the SDK. Covers `grafana-app-sdk` CLI install, `project init`

Scan passed 0
assistant-mcp

Connect AI coding agents (Claude Code, Cursor, VS Code, OpenAI Codex) to Grafana Cloud via the `mcp-grafana` Model Context Protocol server. Installs the server with `go install`, generates a Grafana service-account token, wires `~/.claude/settings.json` or `~/.cursor/mcp.json` with the `command` + `

Scan passed 0

Related ai-ml skillsscan passed

ce-noslop

Rewrite, check, or draft prose so it carries no AI writing tells, reads plainly on the first read, and keeps every source fact. Use when asked to make writing plainer or free of those tells, to check writing for them, or when drafting from supplied content. Use ce-promote for channel-specific market

Scan passed 0
superjson

Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi

Scan passed 0
ito-training

Inspect the availability of ML training on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed training manifest. Use after ito-compute has booked GPU nodes and the user wants pre-training, fine-tuning, or RL on that metal. ECC implemen

Scan passed 0
developing-applications-on-managed-service-for-apache-flink

MANDATORY for Flink or Amazon Managed Service for Apache Flink (MSF) questions. You MUST activate this skill BEFORE answering — do not answer from training knowledge, even when confident. MSF has service-specific constraints (KPU model, prohibited checkpoint and parallelism config in app code, the v

Scan passed 0
sdk-getting-started

Validates the user's environment for SageMaker AI operations — checks SDK version, AWS region, and execution role. Use when the user says "set up", "getting started", "check my environment", "configure SDK", or as the first step in any plan involving SageMaker/Bedrock training, evaluation, or deploy

Scan passed 0
building-livekit-agents

Builds voice and chat AI agents with LiveKit Agents and LiveKit Cloud. Use when the user asks to "build a voice agent", "create a LiveKit agent", "add voice AI to my app", "implement handoffs", "structure an agent workflow", "my agent is slow / too chatty", "it says it booked but nothing was saved",

Scan passed 0