agent-platform-tuning-management
Manages GenAI tuning jobs in Agent Platform. Use this to list, get, or cancel ongoing model tuning jobs. Don't use for fine-tuning models (use `agent-platform-tuning`), deploying models to endpoints (use `agent-platform-deploy`), or managing serving endpoints (use `agent-platform-endpoint-management
- 0
- Installs
- —
- Rating
- —
- Success rate
- 1
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 bb2ea34f8442591c… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Agent Platform Tuning Management
This skill provides instructions on how to manage GenAI Tuning Jobs using the Agent Platform Python SDK. Use this skill when a user wants to check the status of their tuning runs, find an active tuning job, or cancel a job that is running too long.
Safety & Confirmation Tiers (CRITICAL)
Before executing any commands on behalf of the user, you MUST adhere to the following safety tiers based on the action requested:
- Tier R: Read-only (
list,get)- Rule: No confirmation needed. You may execute these commands immediately to gather information for the user.
- Tier D: Destructive & Interruptive (
cancel)- Rule: Cancellation is a Tier D action requiring explicit typed confirmation (e.g. "I confirm" or "Yes, cancel it").
- Required Fields in Dry-Run Confirmation Card: Before cancelling a
tuning job, you MUST present a dry-run confirmation preview clearly
listing:
- Target Resource: The full tuning job resource name or ID (e.g.
projects/<PROJECT_ID>/locations/<REGION>/tuningJobs/<JOB_ID>). - Command / Script: The exact cancellation command or Python code to be executed.
- Expected Effect: Stops the ongoing tuning job; any in-progress training will be halted and cannot be resumed.
- Ask the user to explicitly confirm (e.g., "Do you confirm? Please reply with 'I confirm' or 'Yes, cancel it'.").
- Target Resource: The full tuning job resource name or ID (e.g.
- Same-turn restriction: NEVER execute the cancellation in the same turn as presenting the preview card. Stop immediately and wait for the user to confirm in a new turn. Even if the user provided pre-emptive confirmation (e.g. "Yes, I confirm, cancel tuning job ...") or provides a corrected job ID, you MUST present the dry-run preview for that specific job ID and wait for confirmation in a separate turn before issuing the cancellation.
Phase 0: Environment Setup
CRITICAL: Before running any of the Python snippets below, you MUST ensure the environment is correctly initialized by following these steps:
-
Google Cloud Authentication: Authenticate with your Google Cloud account and configure active Application Default Credentials (ADC) for Agent Platform access:
gcloud auth login gcloud auth application-default login -
Python Dependencies: This skill needs
google-cloud-aiplatform. Do not create a virtual environment — it starts empty and hides packages the environment already provides, forcing a redundant install. Probe, and install only what is missing:python3 -c "import vertexai" || pip install google-cloud-aiplatform -
Execution: Run Python snippets with a plain
python3. There is no environment to activate first.
Workflow Decision Tree
-
Information Gathering: Do you have a Project ID and Region?
- No -> You MUST ask the user for the missing Project ID and Region in plain text, or advise them to check their gcloud configuration. If neither location has this information, then ask the user to provide it. Do not attempt to search random regions on your own.
- Yes -> Proceed to Step 2.
-
Task Type: What does the user want to do?
- Find or List Jobs -> Use the Python SDK to list tuning jobs. (Tier R)
- Check Status / Inspect a Specific Job -> Use the Python SDK to get tuning job details. (Tier R)
- Cancel a Job -> Ask for confirmation, then use the Python SDK to cancel the tuning job. (Tier D)
Using the Python SDK
[!NOTE]
Resource Verification & Missing Projects/Jobs: If the execution of the Python snippet fails with an error (such as
403 Permission Denied,404 Not Found,INVALID_ARGUMENT, or indicating a dummy/missing project or job ID), you MUST inform the user that the project or tuning job does not exist or cannot be accessed. You MUST prompt the user to provide a valid Project ID or Job ID, and stop tool execution immediately to wait for their response. Do NOT retry or loop, do NOT assume the resource is valid, and do NOT execute further scripts before receiving valid details from the user.
1. Listing Tuning Jobs (Tier R)
If the user asks "What tuning jobs do I have running?" or wants to find a specific job ID:
from google.cloud import aiplatform_v1
project_id = "YOUR_PROJECT_ID"
region = "YOUR_REGION"
parent = f"projects/{project_id}/locations/{region}"
client = aiplatform_v1.GenAiTuningServiceClient(
client_options={"api_endpoint": f"{region}-aiplatform.googleapis.com"}
)
jobs = client.list_tuning_jobs(parent=parent)
for job in jobs:
print(f"Name: {job.name}")
print(f"Base Model: {job.base_model}")
print(f"State: {job.state}")
list_tuning_jobs takes no page_size keyword. To set one, pass the request:
client.list_tuning_jobs(request={"parent": parent, "page_size": 50}).
2. Getting Details for a Specific Job (Tier R)
If the user provides a Tuning Job ID and asks for its status:
from google.cloud import aiplatform_v1
project_id = "YOUR_PROJECT_ID"
region = "YOUR_REGION"
job_id = "YOUR_JOB_ID" # 19-digit ID
name = f"projects/{project_id}/locations/{region}/tuningJobs/{job_id}"
client = aiplatform_v1.GenAiTuningServiceClient(
client_options={"api_endpoint": f"{region}-aiplatform.googleapis.com"}
)
job = client.get_tuning_job(name=name)
print(f"Name: {job.name}")
print(f"Base Model: {job.base_model}")
print(f"State: {job.state}")
print(f"Tuning Model: {job.tuned_model_display_name}")
3. Canceling a Job (Tier D)
If the user explicitly requests to stop, abort, or cancel a running tuning job:
Safety Check: Action requires explicit typed confirmation before proceeding. You MUST present a dry-run confirmation card listing the Target Resource, Command/Script, and Expected Effect, and ask the user to type "I confirm" or "Yes, cancel it". Even if the user provided confirming language pre-emptively or is providing a corrected/new job ID, you MUST present the preview card for that specific job ID and wait for their explicit approval in a new turn.
[!IMPORTANT]
NEVER pre-emptively execute any cancellation code or command before receiving the user's response in a new turn. You must never speculate or assume that confirmation will be given. Executing cancellation in the same turn as presenting the preview card is a severe safety violation.
from google.cloud import aiplatform_v1
project_id = "YOUR_PROJECT_ID"
region = "YOUR_REGION"
job_id = "YOUR_JOB_ID" # 19-digit ID
name = f"projects/{project_id}/locations/{region}/tuningJobs/{job_id}"
client = aiplatform_v1.GenAiTuningServiceClient(
client_options={"api_endpoint": f"{region}-aiplatform.googleapis.com"}
)
client.cancel_tuning_job(name=name)
print(f"Successfully requested cancellation for {name}")
Files
1- SKILL.md
97f1b77c1d7.4 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from google/skills8
Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st
Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model
Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run
Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be
Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c
Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry
Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.
Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.
Related devops skillsscan passed
Deployment workflows, CI/CD pipeline patterns, Docker containerization, health checks, rollback strategies, and production readiness checklists for web applications. Use when setting up CI/CD, containerizing an app, or checking production readiness before a release.
Configure deployment settings for /land-and-deploy.
Build or maintain Cloudflare Sandbox apps on the stable @cloudflare/sandbox package. Use sandbox-next for preview apps and sandbox-migrate-to-next for stable-to-preview migrations.
Deploy tRPC on AWS Lambda with awsLambdaRequestHandler() from @trpc/server/adapters/aws-lambda for API Gateway v1 (REST, APIGatewayProxyEvent) and v2 (HTTP, APIGatewayProxyEventV2), and Lambda Function URLs. Enable response streaming with awsLambdaStreamingRequestHandler() wrapped in awslambda.strea
Instruments code so production behavior is visible and diagnosable. Use when adding logging, metrics, tracing, or alerting. Use when shipping any feature that runs in production and you need evidence it works. Use when production issues are reported but you can't tell what happened from the availabl
Deploys and manages full-stack web applications (Next.js, Angular) with Server-Side Rendering (SSR) using Firebase App Hosting. Use when deploying Next.js/Angular apps, configuring apphosting.yaml or firebase.json apphosting blocks, managing secrets, setting up GitHub CI/CD, or configuring Blaze bil