google-agents-cli-observability
This skill should be used when the user wants to "set up tracing", "monitor my agent", "configure logging", "add observability", "debug production traffic", or needs guidance on monitoring deployed agents, including ADK (Agent Development Kit) agents. Covers Cloud Trace, prompt-response logging, Big
- 0
- Installs
- —
- Rating
- —
- Success rate
- 5
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 a319e01c150d5ebc… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Observability Guide
Cloud Trace works out of the box — no infrastructure needed. Prompt-response logging and BigQuery Agent Analytics require Terraform-provisioned infrastructure (service account, GCS bucket, BigQuery dataset). Run
agents-cli infra single-project --apply --project PROJECT_IDto provision these resources. Go projects get the BigQuery telemetry stack too; the GCS completion upload behind prompt-response logging and the BigQuery Agent Analytics plugin are Python only. Seereferences/cloud-trace-and-logging.mdfor details, env vars, and verification commands. If your project isn't scaffolded yet, see/google-agents-cli-scaffoldfirst.
Deployment order
| Do you want Terraform-managed observability? | Action |
|---|---|
| No | Run agents-cli deploy directly (works for all targets). |
| Yes | Run agents-cli infra single-project --apply first, then agents-cli deploy. |
If you already ran agents-cli deploy imperatively:
Do not apply Terraform afterward. For Agent Runtime, Terraform owns the whole engine (service account, deployment spec, env vars), so a later apply can't be reconciled without taking over the resource.
Either delete the deployment and start over, or keep the SDK-deployed instance — skip infra single-project and set the observability env vars by re-running agents-cli deploy --update-env-vars "KEY=VALUE,..."; deploy matches the existing Reasoning Engine by display name and updates it in place, preserving env vars set outside the deploy. You must also grant its service account the telemetry IAM roles the Terraform module would otherwise provision: roles/storage.admin (write completions to the logs bucket), roles/logging.logWriter, roles/cloudtrace.agent, plus roles/bigquery.dataOwner + roles/bigquery.jobUser when scaffolded with --bq-analytics. The full set lives in deployment/terraform/single-project/iam.tf (from app_sa_roles) and telemetry.tf. Terraform-managed env vars aren't available in this mode.
Reference Files
| File | Contents |
|---|---|
references/cloud-trace-and-logging.md | Scaffolded project details — Terraform-provisioned resources, environment variables, verification commands, enabling/disabling locally |
references/bigquery-agent-analytics.md | BQ Agent Analytics plugin — enabling, key features, GCS offloading, tool provenance |
references/adk-docs.md | ADK: adk.dev pages to fetch for detail beyond this skill |
references/feedback-mechanism.md | Adding a user-feedback endpoint — request model, structured logging, log sink → BigQuery |
Observability Tiers
Choose the right level of observability based on your needs:
| Tier | What It Does | Scope | Default State | Best For |
|---|---|---|---|---|
| Cloud Trace | Distributed tracing — execution flow, latency, errors via OpenTelemetry spans | All templates, all environments | Always enabled | Debugging latency, understanding agent execution flow |
| Prompt-Response Logging | GenAI interactions exported to GCS, BigQuery, and Cloud Logging | Scaffolded ADK Python projects | Disabled locally, enabled when deployed | Auditing LLM interactions, compliance |
| BigQuery Agent Analytics | Structured agent events (LLM calls, tool use, outcomes) to BigQuery | ADK Python agents with the plugin enabled | Opt-in (--bq-analytics at scaffold time) | Conversational analytics, custom dashboards, LLM-as-judge evals |
| Third-Party Integrations | External observability platforms (AgentOps, Phoenix, MLflow, etc.) | Any OpenTelemetry-instrumented agent | Opt-in, per-provider setup | Team collaboration, specialized visualization, prompt management |
Ask the user which tier(s) they need — they can be combined. Cloud Trace is always on; the others are additive.
Cloud Trace
Scaffolded agents use OpenTelemetry to emit distributed traces. Every agent invocation produces spans that track the full execution flow.
Span Hierarchy
ADK projects. These are ADK's span names; other frameworks emit their own (
generate_contentcomes from the shared google-genai instrumentor either way).
invoke_workflow (top-level run)
└── invoke_agent (one per agent in the chain)
├── call_llm (model request)
│ └── generate_content (underlying GenAI model call)
└── execute_tool (tool execution)
Setup by Deployment Type
| Deployment | Setup |
|---|---|
| Agent Runtime | Automatic — exporters wired at startup, gated on GOOGLE_CLOUD_AGENT_ENGINE_ENABLE_TELEMETRY (set by deploy); exports to Cloud Trace/Logging + Agent Engine console |
| Cloud Run / GKE (scaffolded) | Automatic — exporters wired at startup, exports to Cloud Trace/Logging |
| Cloud Run / GKE (manual) | Configure OpenTelemetry exporter in your app |
| Local dev | Works with agents-cli playground; traces visible in Cloud Console |
Wired at app startup — ADK Python: get_fast_api_app(otel_to_cloud=True) in app/fast_api_app.py; ADK Go: setupObservability() in observability.go; other templates call their own. references/cloud-trace-and-logging.md has the details.
View traces: Cloud Console → Trace → Trace explorer
ADK: for detailed setup instructions (Agent Runtime CLI/SDK, Cloud Run, custom deployments), fetch https://adk.dev/integrations/cloud-trace/index.md.
Prompt-Response Logging
Captures GenAI interactions and exports to GCS (JSONL) and BigQuery (via log sinks + external tables). Content is governed by two independent tiers; the net Terraform-deploy default is full content in GCS/BigQuery, none in traces:
| Tier | Captures | Controlled by | Default (Terraform deploy) |
|---|---|---|---|
| GCS/BigQuery completions | Full prompts/responses (the prompt-response logging feature) | OTEL_INSTRUMENTATION_GENAI_COMPLETION_HOOK=upload + LOGS_BUCKET_NAME | On — full content |
| Trace spans / Cloud Logging events | Span/event content | OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT (plus ADK_CAPTURE_MESSAGE_CONTENT_IN_SPANS=false, ADK Python only) | Off — NO_CONTENT |
The tiers are independent: GCS/BigQuery uploads capture full content whenever their upload vars are set and do not honor OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT, which governs the traces/events tier only.
ADK Python reads it as the experimental-semconv enum:
NO_CONTENT— no content in spans/events (scaffolded default)EVENT_ONLY— content in Cloud Logging eventsSPAN_ONLY/SPAN_AND_EVENT— content in trace spanstrue/false— invalid; fall back toNO_CONTENT
ADK Go reads the same variable as a boolean: "1" or "true" capture content, every other
value — including the enum members above — elides it.
For the full mechanics (semconv opt-in, declarative Terraform config, env-var table, enabling/disabling, verification commands), see references/cloud-trace-and-logging.md. For ADK logging docs (log levels, configuration, debugging), fetch https://adk.dev/observability/logging/index.md.
BigQuery Agent Analytics Plugin
ADK projects. Optional ADK plugin that logs structured agent events to BigQuery. Enable with
--bq-analyticsat scaffold time. Seereferences/bigquery-agent-analytics.mdfor details.
Third-Party Integrations
Many third-party observability platforms can ingest agent telemetry (via OpenTelemetry or custom instrumentation). The table below covers common ones; the full list is larger (see the pointer below it).
| Platform | Key Differentiator | Setup Complexity | Self-Hosted Option |
|---|---|---|---|
| AgentOps | Session replays, 2-line setup, replaces native telemetry | Minimal | No (SaaS) |
| Arize AX | Commercial platform, production monitoring, evaluation dashboards | Low | No (SaaS) |
| Phoenix | Open-source, custom evaluators, experiment testing | Low | Yes |
| MLflow | OTel traces to MLflow Tracking Server, span tree visualization | Medium (needs SQL backend) | Yes |
| Monocle | 1-call setup, VS Code Gantt chart visualizer | Minimal | Yes (local files) |
| Weave | W&B platform, team collaboration, timeline views | Low | No (SaaS) |
| Freeplay | Prompt management + evals + observability in one platform | Low | No (SaaS) |
Ask the user which platform they prefer — present the trade-offs and let them choose. ADK: fetch a platform's setup page at https://adk.dev/integrations/<slug>/index.md (slugs for the table above: agentops, arize-ax, phoenix, mlflow-tracing, monocle, weave, freeplay); ADK has more observability integrations (Datadog, Galileo, LangWatch, Latitude, Future AGI, Respan, Zespan, …) — browse the complete, current list at https://adk.dev/integrations/ (observability topic). On other frameworks the OpenTelemetry-based platforms still work, but follow the platform's own setup docs.
Troubleshooting
| Issue | Solution |
|---|---|
| No traces in Cloud Trace | Verify telemetry setup runs at startup and the SA has the cloudtrace.agent role. ADK Python: fast_api_app.py uses get_fast_api_app(otel_to_cloud=True); ADK Go: observability.go build the exporter manually. Agent Runtime additionally gates this on GOOGLE_CLOUD_AGENT_ENGINE_ENABLE_TELEMETRY. |
| Prompt-response data not appearing | Check LOGS_BUCKET_NAME is set; verify SA has storage.objectCreator on the bucket; check app logs for telemetry setup warnings |
| Content in traces/events (unwanted) | ADK Python: OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=NO_CONTENT keeps content out of spans/events. ADK Go: any value other than 1/true does, so unset it or set false. NOTE: GCS/BigQuery completions still capture full content — to stop that, remove LOGS_BUCKET_NAME/OTEL_INSTRUMENTATION_GENAI_COMPLETION_HOOK (drop the upload block in service.tf) |
| BigQuery Analytics not logging | ADK Python: verify the plugin is configured in app/agent.py; check BQ_ANALYTICS_DATASET_ID env var is set |
| Third-party integration not capturing spans | Check provider-specific env vars (API keys, endpoints); some providers (AgentOps) replace native telemetry |
| Traces missing tool spans | ADK: tool execution spans appear under execute_tool (other frameworks use their own span names) — check trace explorer filters |
| High telemetry costs | Turn content capture off (NO_CONTENT in Python, false in Go); reduce BigQuery retention; disable unused tiers |
Related Skills
/google-agents-cli-deploy— Deployment targets, CI/CD pipelines, and production workflows/google-agents-cli-workflow— Development workflow, coding guidelines, and operational rules/google-agents-cli-adk-code— ADK API quick reference for writing agent code, Python and Go
Files
5- SKILL.md
4f3638ed9311.6 KB - references/adk-docs.md
d876b98b50445 B - references/bigquery-agent-analytics.md
8bc6ca6c291.3 KB - references/cloud-trace-and-logging.md
5b931ebcc88.2 KB - references/feedback-mechanism.md
021074907d3.9 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from google/agents-cli6
This skill should be used when the user wants to "write agent code", "build an agent with ADK", "add a tool", "create a callback", "define an agent", "use state management" — in a project that needs ADK (Agent Development Kit) API patterns and code examples. It provides a quick reference for agent t
This skill should be used when the user wants to "deploy an agent", "deploy my ADK agent", "set up CI/CD", "configure secrets", "troubleshoot a deployment", or needs guidance on Agent Runtime, Cloud Run, or GKE deployment targets, or binding an agent to an Agent Gateway. Covers deployment workflows,
This skill should be used when the user wants to "run an evaluation", "evaluate my agent", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs guidance on the Agent Platform eval methodology and the Quality Flywheel. Covers ev
This skill should be used when the user wants to "publish an agent", "publish my ADK agent", "register an agent with Gemini Enterprise", "publish to Gemini Enterprise", or needs guidance on the agents-cli publish gemini-enterprise command. Also use when the user wants to "manage agents in Agent Regi
This skill should be used when the user wants to "create an agent project", "start a new ADK project", "build me a new agent", "add CI/CD to my project", "add deployment", "enhance my project", or "upgrade my project". Part of the agents-cli skills suite. Covers `agents-cli scaffold create`, `scaffo
This skill should be used when the user wants to "develop an agent", "build an agent using ADK", "run the agent locally", "debug agent code", "test an agent", "deploy an agent", "publish an agent", "monitor an agent", or needs the ADK (Agent Development Kit) development lifecycle and coding guidelin
Related devops skillsscan passed
Deployment workflows, CI/CD pipeline patterns, Docker containerization, health checks, rollback strategies, and production readiness checklists for web applications. Use when setting up CI/CD, containerizing an app, or checking production readiness before a release.
Configure deployment settings for /land-and-deploy.
Build, migrate, and deploy Next.js apps on Cloudflare Workers with vinext. Use when starting a Next.js project on Cloudflare, moving an existing app to Workers, choosing between vinext and OpenNext, or setting up vinext for Workers. For setup, migration, or deployment, install vinext's upstream skil
Deploy tRPC on AWS Lambda with awsLambdaRequestHandler() from @trpc/server/adapters/aws-lambda for API Gateway v1 (REST, APIGatewayProxyEvent) and v2 (HTTP, APIGatewayProxyEventV2), and Lambda Function URLs. Enable response streaming with awsLambdaStreamingRequestHandler() wrapped in awslambda.strea
Instruments code so production behavior is visible and diagnosable. Use when adding logging, metrics, tracing, or alerting. Use when shipping any feature that runs in production and you need evidence it works. Use when production issues are reported but you can't tell what happened from the availabl
Deploys and manages full-stack web applications (Next.js, Angular) with Server-Side Rendering (SSR) using Firebase App Hosting. Use when deploying Next.js/Angular apps, configuring apphosting.yaml or firebase.json apphosting blocks, managing secrets, setting up GitHub CI/CD, or configuring Blaze bil