skills/ google/skills

cloud-run-alert-configuration

Configures best-practice, high-signal alerting policies for Cloud Run resources on Google Cloud (services, jobs, and worker pools) based on seasoned SRE practices. Use when analyzing, recommending, writing, or deploying Terraform PromQL alerting policies to monitor Cloud Run error rates (4xx/5xx), r

0
Installs
—
Rating
—
Success rate
4
Files scanned
Scan passeddevops
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

4 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 a7eb8ce72926792c… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Cloud Run Alert Configuration

Production-grade observability for Cloud Run using Terraform and PromQL (Cloud Monitoring). Grounded in SRE practices, this skill focuses strictly on actionable user impact and scaling bounds.


CRITICAL RULES

  • Prompt-First Fast Path (Skip Discovery When Named):
    • If the user prompt explicitly specifies the target Cloud Run service, job, or worker pool name (e.g., 'video-encoder', 'nightly-reconciliation', 'web-frontend', 'catalog-service', 'api-gateway', 'order-processor'), SKIP all workspace .tf file scanning (find_by_name, code_search, list_dir) and gcloud CLI discovery commands entirely.
    • Do NOT run gcloud, terraform, or file search tools when the target name is already provided in the prompt. Instead, parameterize the project ID (variable "scoping_project_id" { default = "my-gcp-project" }) and target resource name in Terraform variables and proceed immediately to Step 2 (Configure Alerts).
  • Autonomous Discovery (Only When Target Name is Omitted):
    • Never Scan Root Monorepo or Unbounded Directories: Never run find_by_name or ls across root workspace directories.
    • Config First: Only if the prompt omits the resource name, check .tf files in the immediate working directory for google_cloud_run_v2_service, google_cloud_run_service, or google_cloud_run_v2_job.
    • CLI Second (Graceful Fallback): Only if unconfigured in prompt or local .tf files, attempt gcloud config get-value project and gcloud run services list. If any gcloud command fails (e.g., auth or metadata errors) or terraform is missing, immediately stop running CLI commands and output parameterized HCL using explicit variable defaults.
  • Workload Routing: Always classify the workload target and follow its specific reference guide:
  • Explicit Defaults & User Overrides:
    • Always use explicit defaults for all constants specified in the target workload's reference file (SLO targets, latency thresholds, SLAs, saturation ceilings, rate guards).
    • State the defaults being applied in the final summary output and clearly notify the user that any default constant can be customized or overridden via Terraform variables or prompt input.
  • Metric Scope Centralization: Parameterize project = var.scoping_project_id in all Terraform google_monitoring_alert_policy resources so the policy can target either a single project or a centralized Cloud Monitoring Metrics Scope.
  • PromQL duration (Retest Window) Rules:
    • Lookbacks $\le$ 25h: Set duration = "300s" (5m buffer) to absorb transient blips and scale-up lag (except immediate job failure alerts which use duration = "0s").
    • Lookbacks $> 25$h (e.g. 3d/7d Slow Burn): Omit duration entirely (or set to 0s). Cloud Monitoring rejects PromQL queries with duration set on lookbacks >25h (INVALID_ARGUMENT).
  • Terraform Standards & Mandatory Labels:
    • Output clean, complete .tf configurations using google_monitoring_alert_policy and condition_prometheus_query_language directly in your response.
    • Mandatory User Labels: Every google_monitoring_alert_policy resource MUST include a user_labels block containing:
      user_labels = {
        created-with-google-skill = "cloud-run-alert-configuration"
      }
      
    • Include alert_strategy { auto_close = "604800s" } and parameterize notification_channels = var.notification_channels.

WORKFLOW STEPS

1. Discovery & Target Identification

  • Fast Path (Target Named in Prompt): If the user prompt names the target Cloud Run service, job, or worker pool, skip all discovery commands and file searches and proceed directly to Step 2.
  • Discovery Fallback (Target Unnamed): Only if no resource name is provided in the prompt, check local .tf files or run gcloud to identify the target workload type and name. If gcloud auth fails, fall back immediately to default Terraform variables (var.scoping_project_id).

2. Configure Alerts

  • Route to the corresponding guide to generate the alert policies:
    • HTTP Services: Open services.md. Apply the requested alerting policy or standard suite covering availability SLOs (5xx), request latency (P95/P99), client errors (4xx), container instance saturation, container CPU/memory utilization, traffic anomalies (drop/surge), and billable instance time.
    • Batch Jobs: Open jobs.md. Apply immediate job execution failure alerts (duration = "0s").
    • Worker Pools: Open worker_pools.md. Apply the 4-policy standard suite (Task Success SLO Fast/Slow Burn, Backlog ETD, Message Age SLA).

3. Terraform Generation & Review

  • Provide the complete HCL configuration in your response with explicitly parameterized defaults and the mandatory user_labels block (created-with-google-skill = "cloud-run-alert-configuration").
  • State the applied defaults and remind the user of their ability to override any constant.
  • Provide a clear plain-English breakdown of the PromQL logic and triggering thresholds.

Additional Resources

Files

4
50.1 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from google/skills8

agent-platform-alert-configuration

Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st

Needs review 0
agent-platform-deploy

Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model

Scan passed 0
agent-platform-endpoint-management

Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run

Scan passed 0
agent-platform-eval-flywheel

Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be

Scan passed 0
agent-platform-inference

Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c

Scan passed 0
agent-platform-migrate-from-ai-studio

Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry

Scan passed 0
agent-platform-model-registry

Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.

Scan passed 0
agent-platform-prompt-management

Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.

Scan passed 0

Related devops skillsscan passed

deployment-patterns

Deployment workflows, CI/CD pipeline patterns, Docker containerization, health checks, rollback strategies, and production readiness checklists for web applications. Use when setting up CI/CD, containerizing an app, or checking production readiness before a release.

Scan passed 0
setup-deploy

Configure deployment settings for /land-and-deploy.

Scan passed 0
sandbox-stable

Build or maintain Cloudflare Sandbox apps on the stable @cloudflare/sandbox package. Use sandbox-next for preview apps and sandbox-migrate-to-next for stable-to-preview migrations.

Scan passed 0
adapter-aws-lambda

Deploy tRPC on AWS Lambda with awsLambdaRequestHandler() from @trpc/server/adapters/aws-lambda for API Gateway v1 (REST, APIGatewayProxyEvent) and v2 (HTTP, APIGatewayProxyEventV2), and Lambda Function URLs. Enable response streaming with awsLambdaStreamingRequestHandler() wrapped in awslambda.strea

Scan passed 0
observability-and-instrumentation

Instruments code so production behavior is visible and diagnosable. Use when adding logging, metrics, tracing, or alerting. Use when shipping any feature that runs in production and you need evidence it works. Use when production issues are reported but you can't tell what happened from the availabl

Scan passed 0
firebase-app-hosting-basics

Deploys and manages full-stack web applications (Next.js, Angular) with Server-Side Rendering (SSR) using Firebase App Hosting. Use when deploying Next.js/Angular apps, configuring apphosting.yaml or firebase.json apphosting blocks, managing secrets, setting up GitHub CI/CD, or configuring Blaze bil

Scan passed 0