skills/ google/skills

cloud-monitoring-promql-query

Generates valid PromQL queries from Cloud Monitoring metric descriptors and resource parameters. Use when asked to create, generate, write, or format PromQL queries, PromQL strings, or PromQL aggregations for Cloud Monitoring metrics and resources. Don't use for raw metric discovery or metric select

0
Installs
—
Rating
—
Success rate
5
Files scanned
Scan passeddevops
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

5 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 47f12a239fcba1db… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Cloud Monitoring PromQL Generator

Use this skill to generate a valid PromQL query from any Cloud Monitoring metric type. This guide applies to all Cloud Monitoring metric types by mapping Cloud Monitoring metric and resource descriptors to PromQL structures.

Workflow

Resolve Project ID (CRITICAL & BLOCKING)

Before performing any other actions (such as searching code, reading references, or running validation), you MUST verify whether the Google Cloud Project ID is available:

  1. Check Prompt/Payload: Look for the Project ID in the user's prompt or input.
  2. Check Environment: If the Project ID is not present in the prompt, you MUST run gcloud config get-value project to attempt to resolve it from the environment.
  3. Ask for Clarification (BLOCKING): If the Project ID is not in the prompt AND the gcloud command fails, returns an empty string, or is unavailable, you MUST immediately stop. Do NOT generate a PromQL query, do not run the validation script, and do not use placeholders (like YOUR_PROJECT_ID). You must refuse to proceed and ask the user to provide the Project ID.

Inspect Metric and Resource Descriptors

  1. Use Provided Descriptors First: If the user's prompt already includes metric descriptor details (such as metric.type, metricKind, valueType, or monitoredResourceTypes) or specific resource filter values, use those values directly instead of calling the Cloud Monitoring API.
  2. Discover Missing Descriptors: If exact metric descriptors (metric.type, metricKind, valueType) are missing or underspecified, resolve the target metric type's descriptor using one of these paths:
    • Vague Query: If the prompt is vague (for example, "VM CPU usage"), use the cloud-monitoring-metric-selection skill first to identify the specific metric type.
    • Known Metric Type: If you already have the specific metric type name (for example, compute.googleapis.com/instance/cpu/utilization) but need its descriptor, call the google-cloud-monitoring:list_metric_descriptors MCP tool. If the tool is missing, refer to the cloud-monitoring-metric-selection skill to configure the Cloud Monitoring MCP server.
    • Fallback: If the MCP tool cannot be configured, fall back to making a direct Cloud Monitoring API call.
  3. Identify Key Fields: From the retrieved descriptor, identify four key schema attributes:
    • type: The Cloud Monitoring metric type string.
    • metricKind: GAUGE, DELTA, or CUMULATIVE.
    • valueType: INT64, DOUBLE, DISTRIBUTION, or BOOL.
    • monitoredResourceTypes: Compatible resource.type strings required for resource scoping and grouping.

Resolve Resource Filters & Discovery Protocol

To filter data by a specific resource instance, apply these resource rules and discovery protocols:

  1. Monitored Resource Filter: Always include the monitored_resource="<type>" filter in your query to prevent collisions across services that share metric names.

    • Example: monitored_resource="gae_app"
  2. Preserve User Literals (CRITICAL): ALWAYS use the literal resource names, namespaces, and IDs provided in the user's prompt. Do NOT override or replace these values with active resource names found during Cloud Monitoring discovery unless the user explicitly asked you to find active resources. Telemetry discovery must only be used to identify metric type names and label keys, not to override user input.

  3. Resource Identifier Mapping:

    • Direct & Specific Keys: Use the most specific resource identifier available. Example: version_id, cluster_name.
    • Name-to-ID Resolution: If the user filters by a resource name (such as "instance-1"), but the resource schema uses numeric IDs (like instance_id), use PromQL string name labels instead of numeric ID labels. Example: instance_name, metadata_system_name.
    • Composite Identifiers: For resources with hierarchical identifiers (such as Cloud SQL databases), format the filter as a single composite key. Do NOT split them into separate project_id and sub-resource labels. Example: database_id="{project_id}:{instance_name}".
  4. Resource Label Discovery: The google-cloud-monitoring:list_metric_descriptors tool only returns metric-specific labels. If the label schema for a monitored resource is unknown, fetch the resource descriptor directly from the Cloud Monitoring v3 REST API (projects.monitoredResourceDescriptors.get):

    TOKEN=$(gcloud auth application-default print-access-token 2>/dev/null || gcloud auth print-access-token)
    curl -s -H "Authorization: Bearer ${TOKEN}" \
    "https://monitoring.googleapis.com/v3/projects/{project_id}/monitoredResourceDescriptors/{monitored_resource_type}"
    

    An HTTP 200 OK response returns the MonitoredResourceDescriptor object containing the labels array with the exact resource label keys for that resource.

Choose Aggregation Structure & Defaults

The query structure and aggregation functions (such as rate, histogram_quantile, sum, or avg) depend on the metric type and how it is visualized.

  1. Consult the Reference: Consult the Cloud Monitoring to PromQL Basic Aggregations Reference as the single source of truth to map Cloud Monitoring properties (Metric Kind, Value Type, Aligner, Reducer) to their PromQL structures.
  2. SRE Aggregation & Visualization Rules:
    • Do NOT sum or average ratio/percentage utilization metrics (like CPU % or Memory limit utilization) across resource instances. Instead, keep them unaggregated (raw metric), group by instance, or wrap in topk(30, avg_over_time(...)).
    • State Label Filtering (CRITICAL): Only the metrics agent.googleapis.com/memory/percent_used and agent.googleapis.com/disk/percent_used require {state!="free"}. Do NOT filter by {state="used"}.

Format & Validate Query

Before presenting any PromQL queries, validate them using the linter:

Python Dependencies

Before executing the validation script (scripts/validate_promql.py), install the required Python dependencies:

python3 -c "import promql_parser" || pip install promql-parser

Validation Procedure

  1. Format Constraints:
    • Metric Name Normalization: Convert Cloud Monitoring metric types to PromQL metric names using this recipe:
      1. Split Domain and Path: Split the Cloud Monitoring metric type by the first slash (/) to separate the domain from the path.
        • Example: storage.googleapis.com/network/received_bytes_count -> domain storage.googleapis.com, path network/received_bytes_count
      2. Normalize Domain: Replace all periods (.) in the domain with underscores (_).
        • Example: storage.googleapis.com -> storage_googleapis_com
      3. Normalize Path: Replace all periods (.) and slashes (/) in the path with underscores (_).
        • Example: network/received_bytes_count -> network_received_bytes_count
      4. Join with Colon: Join the normalized domain and normalized path with a colon (:).
        • Example: storage_googleapis_com:network_received_bytes_count
      5. Native Prometheus Metrics: If the metric type has no slash, keep it as-is.
        • Example: up -> up, http_requests_total -> http_requests_total
      6. Distribution Suffix: If the metric's valueType is DISTRIBUTION, append _bucket to the end of the normalized name.
        • Example: cloudfunctions.googleapis.com/function/execution_times -> cloudfunctions_googleapis_com:function_execution_times_bucket
    • Ensure the final query is a single line with no comments (no # or //). Cloud Monitoring query translation collapses whitespace and can cause code trailing a comment to be ignored or throw parsing errors.
    • Grouping Clause Syntax: Ensure grouping clauses (such as by (label)) only follow aggregation operators (such as sum, avg, min, max, or count). Never place a grouping clause directly after a metric selector.
      • Incorrect: metric{...} by (label)
      • Correct: sum(rate(metric{...}[5m])) by (label)
    • Fenced Output Code Block: ALWAYS wrap the final verified PromQL query in a fenced promql code block in your final response.
  2. Linter Verification:
    • Validate all generated queries in a single batch: python3 <path_to_skill>/scripts/validate_promql.py --query '<q1>' '<q2>'
    • If validation fails, read PromQL Error Recovery Guide to diagnose and fix common type mismatches and syntax errors before repeating the loop.

References

Files

5
45.0 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from google/skills8

agent-platform-alert-configuration

Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st

Needs review 0
agent-platform-deploy

Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model

Scan passed 0
agent-platform-endpoint-management

Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run

Scan passed 0
agent-platform-eval-flywheel

Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be

Scan passed 0
agent-platform-inference

Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c

Scan passed 0
agent-platform-migrate-from-ai-studio

Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry

Scan passed 0
agent-platform-model-registry

Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.

Scan passed 0
agent-platform-prompt-management

Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.

Scan passed 0

Related devops skillsscan passed

deployment-patterns

Deployment workflows, CI/CD pipeline patterns, Docker containerization, health checks, rollback strategies, and production readiness checklists for web applications. Use when setting up CI/CD, containerizing an app, or checking production readiness before a release.

Scan passed 0
setup-deploy

Configure deployment settings for /land-and-deploy.

Scan passed 0
sandbox-stable

Build or maintain Cloudflare Sandbox apps on the stable @cloudflare/sandbox package. Use sandbox-next for preview apps and sandbox-migrate-to-next for stable-to-preview migrations.

Scan passed 0
adapter-aws-lambda

Deploy tRPC on AWS Lambda with awsLambdaRequestHandler() from @trpc/server/adapters/aws-lambda for API Gateway v1 (REST, APIGatewayProxyEvent) and v2 (HTTP, APIGatewayProxyEventV2), and Lambda Function URLs. Enable response streaming with awsLambdaStreamingRequestHandler() wrapped in awslambda.strea

Scan passed 0
observability-and-instrumentation

Instruments code so production behavior is visible and diagnosable. Use when adding logging, metrics, tracing, or alerting. Use when shipping any feature that runs in production and you need evidence it works. Use when production issues are reported but you can't tell what happened from the availabl

Scan passed 0
firebase-app-hosting-basics

Deploys and manages full-stack web applications (Next.js, Angular) with Server-Side Rendering (SSR) using Firebase App Hosting. Use when deploying Next.js/Angular apps, configuring apphosting.yaml or firebase.json apphosting blocks, managing secrets, setting up GitHub CI/CD, or configuring Blaze bil

Scan passed 0