agent-platform-endpoint-management
Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run
- 0
- Installs
- —
- Rating
- —
- Success rate
- 1
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 096177b0b305cb2f… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Agent Platform Endpoint Management
Overview
This skill provides procedural knowledge for managing Agent Platform Endpoints. Endpoints are logical serving hosts that provide a stable URL for online predictions. You must create an endpoint before you can deploy a model to it.
Safety & Confirmation Tiers (CRITICAL)
Before executing any commands on behalf of the user, you MUST adhere to the following safety tiers based on the action requested:
- Tier R: Read-only (
list,describe,get)- No confirmation needed. Execute immediately to gather information.
- Tier M: Mutating & Reversible (
create,update)- Requires interactive confirmation with 'Yes'/'No' options. The
confirmation prompt MUST contain the exact, literal command string with
all required flags (e.g.
--region=us-central1,--display-name="...") — natural-language paraphrases are NOT sufficient. - Same-turn restriction: NEVER execute the command in the same turn as presenting the confirmation prompt. Stop and wait for the user's reply; only execute after explicit 'Yes' / approval.
- Requires interactive confirmation with 'Yes'/'No' options. The
confirmation prompt MUST contain the exact, literal command string with
all required flags (e.g.
- Tier D: Destructive & Irreversible (
delete)- Requires explicit typed confirmation (e.g. "I confirm" or "Yes,
delete it"). Ask for confirmation IMMEDIATELY — before any pre-flight
checks (don't
describefirst, don't check if the endpoint is empty first). - Same-turn restriction: NEVER execute in the same turn as asking for typed confirmation. Wait for the user to reply in a new turn.
- Requires explicit typed confirmation (e.g. "I confirm" or "Yes,
delete it"). Ask for confirmation IMMEDIATELY — before any pre-flight
checks (don't
Phase 0: Environment Setup
CRITICAL: Before running any commands, you MUST ensure the environment is correctly initialized by following these steps:
-
Google Cloud Authentication: Authenticate with your Google Cloud credentials and configure active Application Default Credentials (ADC) for Agent Platform access:
gcloud auth login gcloud auth application-default login -
Set Project: Configure the active project for subsequent commands:
gcloud config set project $PROJECT_ID -
Region: Always specify
--region=$LOCATION_IDon each command below. Do NOT useglobal. Ask the user to specify the region if not provided.
1. Listing Endpoints (Tier R)
Use this command to discover existing endpoints in a specific region and retrieve their IDs. No confirmation is required.
gcloud ai endpoints list \
--region=$LOCATION_ID
(Optional) To bound the result, use --limit=$LIMIT; --page-size=$PAGE_SIZE
controls API chunking only and does NOT limit the total output. Use
--page-token=$PAGE_TOKEN to continue from a previous batch.
[!IMPORTANT]
Always specify the
--region. Do NOT use 'global'. Ask the user to specify if not provided.
2. Describing an Endpoint (Tier R)
Retrieve the full metadata for a specific endpoint. No confirmation is required.
gcloud ai endpoints describe $ENDPOINT_ID \
--region=$LOCATION_ID
The models deployed on an endpoint are its deployedModels in this output. In
the Python SDK they are endpoint.gca_resource.deployed_models;
aiplatform.Endpoint has no deployed_models attribute.
3. Creating an Endpoint (Tier M)
Create a new endpoint resource. The parent resource is the location. Action requires an inline confirmation card before proceeding.
gcloud ai endpoints create \
--region=$LOCATION_ID \
--display-name="my-endpoint"
The command has no --asynchronous flag: it waits for the operation and prints
the new endpoint's resource name, whose last segment is the endpoint ID.
[!IMPORTANT]
You MUST seek interactive confirmation first. Your confirmation prompt MUST show the literal command string. For example:
gcloud ai endpoints create --region=$LOCATION_ID --display-name="my-endpoint"Or the exact flags. Do not execute this command in the same turn as proposing the confirmation.
4. Updating an Endpoint (Tier M)
Update endpoint metadata such as display name or labels. Action requires an inline confirmation card before proceeding.
gcloud ai endpoints update $ENDPOINT_ID \
--region=$LOCATION_ID \
--display-name="new-display-name"
Check if the endpoint exists first by either listing or describing the endpoint.
[!IMPORTANT]
You MUST seek interactive confirmation first. Your confirmation prompt MUST show the literal command string. For example:
gcloud ai endpoints update $ENDPOINT_ID --region=$LOCATION_ID --display-name="new-display-name"Or the exact flags. CRITICAL: You are strictly prohibited from executing this command in the same turn as asking for confirmation. When you ask for confirmation, you MUST stop immediately and wait for the user to reply.
5. Deleting an Endpoint (Tier D)
Permanently delete an endpoint resource. Action requires explicit typed confirmation before proceeding.
gcloud ai endpoints delete $ENDPOINT_ID \
--region=$LOCATION_ID
[!WARNING]
All models must be undeployed from the endpoint before it can be deleted. Do not run
describeuntil AFTER you have received typed confirmation to delete.
6. Traffic Splitting (Tier M)
You can manage traffic split between different models deployed on the same endpoint during an update. Action requires an inline confirmation card before proceeding.
# Example: Deploying a model with a specific traffic split is usually done
# via 'gcloud ai endpoints deploy-model'.
Refer to the agent-platform-deploy skill for instructions on deploying and
undeploying models.
Troubleshooting
- 403 Permission Denied: Ensure
aiplatform.adminorownerrole is assigned. - Quota Exceeded: Verify the region's endpoint quota in the Cloud Console.
- Resource Busy: If a deletion fails, check if models are still being undeployed.
Files
1- SKILL.md
53776ef6226.4 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from google/skills8
Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st
Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model
Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be
Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c
Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry
Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.
Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.
Manage and query Agent Platform RAG Engine Corpora and retrieve grounded contexts using the Google GenAI SDK. Use when listing RAG corpora or files, inspecting a corpus, retrieving contexts, or generating content grounded in a RAG corpus. Do not use for standard database queries (use SQL/Spanner ski
Related devops skillsscan passed
Build monitoring dashboards that answer real operator questions for Grafana, SigNoz, and similar platforms. Use when turning metrics into a working dashboard instead of a vanity board.
Deploys and configures classic Firebase Hosting for static websites, single-page apps (SPAs), and microservices. Use when deploying static sites/SPAs, setting up custom domains, configuring firebase.json hosting settings (redirects, rewrites, headers, multi-site), or managing preview channels. Don't
Use when creating new skills, editing existing skills, or verifying skills work before deployment
Runs SQL queries on CloudWatch Logs data exported as Apache Iceberg tables in S3 Tables. Covers VPC Flow Logs, WAF logs, CloudFront access logs, Route 53 resolver logs, Network Firewall logs, EKS audit logs, Verified Access logs, SES logs, VPC Lattice logs, Step Functions logs, NLB access logs, and