cloud-run-alert-configuration
Configures best-practice, high-signal alerting policies for Cloud Run resources on Google Cloud (services, jobs, and worker pools) based on seasoned SRE practices. Use when analyzing, recommending, writing, or deploying Terraform PromQL alerting policies to monitor Cloud Run error rates (4xx/5xx), r
- 0
- Installs
- —
- Rating
- —
- Success rate
- 4
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 a7eb8ce72926792c… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Cloud Run Alert Configuration
Production-grade observability for Cloud Run using Terraform and PromQL (Cloud Monitoring). Grounded in SRE practices, this skill focuses strictly on actionable user impact and scaling bounds.
CRITICAL RULES
- Prompt-First Fast Path (Skip Discovery When Named):
- If the user prompt explicitly specifies the target Cloud Run service, job, or worker pool name (e.g.,
'video-encoder','nightly-reconciliation','web-frontend','catalog-service','api-gateway','order-processor'), SKIP all workspace.tffile scanning (find_by_name,code_search,list_dir) andgcloudCLI discovery commands entirely. - Do NOT run
gcloud,terraform, or file search tools when the target name is already provided in the prompt. Instead, parameterize the project ID (variable "scoping_project_id" { default = "my-gcp-project" }) and target resource name in Terraform variables and proceed immediately to Step 2 (Configure Alerts).
- If the user prompt explicitly specifies the target Cloud Run service, job, or worker pool name (e.g.,
- Autonomous Discovery (Only When Target Name is Omitted):
- Never Scan Root Monorepo or Unbounded Directories: Never run
find_by_nameorlsacross root workspace directories. - Config First: Only if the prompt omits the resource name, check
.tffiles in the immediate working directory forgoogle_cloud_run_v2_service,google_cloud_run_service, orgoogle_cloud_run_v2_job. - CLI Second (Graceful Fallback): Only if unconfigured in prompt or local
.tffiles, attemptgcloud config get-value projectandgcloud run services list. If anygcloudcommand fails (e.g., auth or metadata errors) orterraformis missing, immediately stop running CLI commands and output parameterized HCL using explicit variable defaults.
- Never Scan Root Monorepo or Unbounded Directories: Never run
- Workload Routing: Always classify the workload target and follow its
specific reference guide:
- For HTTP Services follow services.md
- For Cloud Run Jobs follow jobs.md
- For Worker Pools follow worker_pools.md
- Explicit Defaults & User Overrides:
- Always use explicit defaults for all constants specified in the target workload's reference file (SLO targets, latency thresholds, SLAs, saturation ceilings, rate guards).
- State the defaults being applied in the final summary output and clearly notify the user that any default constant can be customized or overridden via Terraform variables or prompt input.
- Metric Scope Centralization: Parameterize
project = var.scoping_project_idin all Terraformgoogle_monitoring_alert_policyresources so the policy can target either a single project or a centralized Cloud Monitoring Metrics Scope. - PromQL
duration(Retest Window) Rules:- Lookbacks $\le$ 25h: Set
duration = "300s"(5m buffer) to absorb transient blips and scale-up lag (except immediate job failure alerts which useduration = "0s"). - Lookbacks $> 25$h (e.g. 3d/7d Slow Burn): Omit
durationentirely (or set to0s). Cloud Monitoring rejects PromQL queries withdurationset on lookbacks >25h (INVALID_ARGUMENT).
- Lookbacks $\le$ 25h: Set
- Terraform Standards & Mandatory Labels:
- Output clean, complete
.tfconfigurations usinggoogle_monitoring_alert_policyandcondition_prometheus_query_languagedirectly in your response. - Mandatory User Labels: Every
google_monitoring_alert_policyresource MUST include auser_labelsblock containing:user_labels = { created-with-google-skill = "cloud-run-alert-configuration" } - Include
alert_strategy { auto_close = "604800s" }and parameterizenotification_channels = var.notification_channels.
- Output clean, complete
WORKFLOW STEPS
1. Discovery & Target Identification
- Fast Path (Target Named in Prompt): If the user prompt names the target Cloud Run service, job, or worker pool, skip all discovery commands and file searches and proceed directly to Step 2.
- Discovery Fallback (Target Unnamed): Only if no resource name is provided in the prompt, check local
.tffiles or rungcloudto identify the target workload type and name. Ifgcloudauth fails, fall back immediately to default Terraform variables (var.scoping_project_id).
2. Configure Alerts
- Route to the corresponding guide to generate the alert policies:
- HTTP Services: Open services.md. Apply the requested alerting policy or standard suite covering availability SLOs (5xx), request latency (P95/P99), client errors (4xx), container instance saturation, container CPU/memory utilization, traffic anomalies (drop/surge), and billable instance time.
- Batch Jobs: Open jobs.md. Apply immediate job
execution failure alerts (
duration = "0s"). - Worker Pools: Open worker_pools.md. Apply the 4-policy standard suite (Task Success SLO Fast/Slow Burn, Backlog ETD, Message Age SLA).
3. Terraform Generation & Review
- Provide the complete HCL configuration in your response with explicitly parameterized defaults and the mandatory
user_labelsblock (created-with-google-skill = "cloud-run-alert-configuration"). - State the applied defaults and remind the user of their ability to override any constant.
- Provide a clear plain-English breakdown of the PromQL logic and triggering thresholds.
Additional Resources
Files
4- SKILL.md
6fb8a21af06.8 KB - references/jobs.md
a34d0415af3.3 KB - references/services.md
52e51c487c28.0 KB - references/worker_pools.md
77cd377d4012.0 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from google/skills8
Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st
Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model
Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run
Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be
Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c
Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry
Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.
Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.
Related devops skillsscan passed
Deployment workflows, CI/CD pipeline patterns, Docker containerization, health checks, rollback strategies, and production readiness checklists for web applications. Use when setting up CI/CD, containerizing an app, or checking production readiness before a release.
Configure deployment settings for /land-and-deploy.
Build or maintain Cloudflare Sandbox apps on the stable @cloudflare/sandbox package. Use sandbox-next for preview apps and sandbox-migrate-to-next for stable-to-preview migrations.
Deploy tRPC on AWS Lambda with awsLambdaRequestHandler() from @trpc/server/adapters/aws-lambda for API Gateway v1 (REST, APIGatewayProxyEvent) and v2 (HTTP, APIGatewayProxyEventV2), and Lambda Function URLs. Enable response streaming with awsLambdaStreamingRequestHandler() wrapped in awslambda.strea
Instruments code so production behavior is visible and diagnosable. Use when adding logging, metrics, tracing, or alerting. Use when shipping any feature that runs in production and you need evidence it works. Use when production issues are reported but you can't tell what happened from the availabl
Deploys and manages full-stack web applications (Next.js, Angular) with Server-Side Rendering (SSR) using Firebase App Hosting. Use when deploying Next.js/Angular apps, configuring apphosting.yaml or firebase.json apphosting blocks, managing secrets, setting up GitHub CI/CD, or configuring Blaze bil