skills/ bagelhole/devops-security-agent-skills

sre-dashboards

Design and operationalize SRE dashboards that surface reliability, latency, error, saturation, and capacity signals across services. Use when building observability views for SLOs, incident response, and executive reliability reporting.

0
Installs
—
Rating
—
Success rate
1
Files scanned
Scan passeddevops
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

1 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 d55f43f802f2b8d0… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

SRE Dashboards

Build dashboards that help teams detect, triage, and prevent reliability incidents.

When to Use This Skill

Use this skill when:

  • Defining service-level dashboards for production systems
  • Tracking SLO health and error-budget burn
  • Creating incident command-center views
  • Standardizing dashboard patterns across teams

Prerequisites

  • Metrics pipeline (Prometheus, OpenTelemetry, or vendor equivalent)
  • Logs/traces linked to services and environments
  • Agreed service taxonomy (team, service, tier, environment)

Dashboard Architecture

Structure dashboards in layers:

  1. Executive Reliability View: SLO attainment, incident counts, MTTR trends.
  2. Service Health View: RED/USE metrics, dependency health, release markers.
  3. Deep-Dive View: Per-endpoint latency, resource saturation, error categories.

Keep each view answer-oriented:

  • Are customers impacted?
  • What changed?
  • Where is the bottleneck?

Core SRE Panels

Golden Signals

  • Latency: p50/p95/p99 request duration by endpoint
  • Traffic: request throughput and queue depth
  • Errors: 5xx rate, failed jobs, timeout ratio
  • Saturation: CPU, memory, disk I/O, thread/connection pool exhaustion

SLO Panels

  • Current SLI value (rolling windows: 5m, 1h, 24h, 30d)
  • Error-budget remaining (%)
  • Burn-rate panels (fast and slow windows)
  • Multi-window burn alert status

Change Correlation

  • Deployment markers and config-change annotations
  • Feature flag state overlays
  • Upstream/downstream dependency error rates

Example PromQL Snippets

# API error rate (%)
100 * sum(rate(http_requests_total{status=~"5.."}[5m]))
  / sum(rate(http_requests_total[5m]))
# p95 latency by route
histogram_quantile(0.95,
  sum by (le, route) (rate(http_request_duration_seconds_bucket[5m]))
)
# Fast burn rate (5m / 1h)
(
  sum(rate(http_requests_total{status=~"5.."}[5m]))
  / sum(rate(http_requests_total[5m]))
)
/
(
  sum(rate(http_requests_total{status=~"5.."}[1h]))
  / sum(rate(http_requests_total[1h]))
)

Operational Guidelines

  • Use consistent color semantics (green=healthy, yellow=degrading, red=breach)
  • Label units explicitly (ms, req/s, %, cores)
  • Default time windows to incident-friendly ranges (15m, 1h, 6h, 24h)
  • Minimize panel count per dashboard to reduce cognitive load
  • Add runbook links directly in panel descriptions

Troubleshooting

Panel appears flat or empty

  • Verify label cardinality and filters (service, env, region)
  • Confirm scrape/ingest latency is within expected range
  • Check metric rename regressions after instrumentation updates

High cardinality slows dashboards

  • Aggregate by stable dimensions (service, route_group) instead of raw IDs
  • Use recording rules for expensive percentile and ratio queries
  • Split deep-dive dashboards from NOC summary dashboards

Related Skills

Files

1
3.5 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from bagelhole/devops-security-agent-skills8

access-review

Conduct periodic access reviews and certifications. Implement access governance and recertification workflows. Use when managing access compliance.

Scan passed 0
agent-evals

Build automated evaluation suites for AI agents using golden datasets, rubrics, and regression gates. Use when shipping agent features, validating prompt changes, or gating deployments on quality.

Needs review 0
agent-observability

Instrument AI agents with tracing, token metrics, latency, and cost visibility. Use for reliability and debugging.

Scan passed 0
ai-agent-security

Secure AI agents against prompt injection, tool abuse, and data exfiltration with defense-in-depth controls. Use when building, deploying, or hardening agentic AI systems that invoke tools, access data, or interact with production infrastructure.

Flagged 0
ai-coding-agent-guardrails

Secure AI coding agents (Claude Code, Cursor, Codex, Copilot) with permission boundaries, secret protection, code review gates, and safe sandbox configurations for team environments.

Needs review 0
ai-inference-service-mesh

Use service mesh patterns for AI inference traffic management, mTLS, canary releases, policy enforcement, and cross-cluster resilience.

Scan passed 0
ai-pipeline-orchestration

Orchestrate AI/ML pipelines for data ingestion, model training, batch inference, and RAG indexing using Prefect, Airflow, or Dagster. Build reliable, observable, and retriable workflows for production AI systems.

Scan passed 0
ai-red-teaming

Run structured AI red team exercises for jailbreak resistance, data exfiltration risk, harmful output controls, and agent tool abuse resilience.

Needs review 0

Related devops skillsscan passed

network-config-validation

Pre-deployment checks for router and switch configuration, including dangerous commands, duplicate addresses, subnet overlaps, stale references, management-plane risk, and IOS-style security hygiene. Use when reviewing a router or switch configuration before deployment.

Scan passed 0
land-and-deploy

Land and deploy workflow. (gstack)

Scan passed 0
nextjs-on-cloudflare

Build, migrate, and deploy Next.js apps on Cloudflare Workers with vinext. Use when starting a Next.js project on Cloudflare, moving an existing app to Workers, choosing between vinext and OpenNext, or setting up vinext for Workers. For setup, migration, or deployment, install vinext's upstream skil

Scan passed 0
adapter-aws-lambda

Deploy tRPC on AWS Lambda with awsLambdaRequestHandler() from @trpc/server/adapters/aws-lambda for API Gateway v1 (REST, APIGatewayProxyEvent) and v2 (HTTP, APIGatewayProxyEventV2), and Lambda Function URLs. Enable response streaming with awsLambdaStreamingRequestHandler() wrapped in awslambda.strea

Scan passed 0
observability-and-instrumentation

Instruments code so production behavior is visible and diagnosable. Use when adding logging, metrics, tracing, or alerting. Use when shipping any feature that runs in production and you need evidence it works. Use when production issues are reported but you can't tell what happened from the availabl

Scan passed 0
firebase-app-hosting-basics

Deploys and manages full-stack web applications (Next.js, Angular) with Server-Side Rendering (SSR) using Firebase App Hosting. Use when deploying Next.js/Angular apps, configuring apphosting.yaml or firebase.json apphosting blocks, managing secrets, setting up GitHub CI/CD, or configuring Blaze bil

Scan passed 0