skills/ google/skills

gke-productionize

Orchestrates comprehensive production readiness reviews and assessments for GKE clusters and workloads across scalability, security, reliability, observability, backup/DR, and cost optimization. Use when asked to productionize, prepare, assess, audit, or review a GKE cluster or workload before going

0
Installs
—
Rating
—
Success rate
1
Files scanned
Scan passedsecurity
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

1 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 f2bc6596d83bae68… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

GKE Productionize Skill

This skill acts as a high-level orchestrator for preparing a GKE cluster and its workloads for production readiness.

[!IMPORTANT] This is a meta-skill or orchestrator skill. You are expected to invoke and run many other specialized skills listed in this document as part of the overall productionization process. Do not attempt to implement all production readiness features directly within this skill; instead, use this skill to assess the environment and then delegate to the specific skills for each domain.

Scope

This skill is adaptable to:

  • A single application (already on Kubernetes or not).
  • A set of applications.
  • A target cluster.

Workflow

1. Discovery Phase

Before making recommendations, discover the current state of the environment.

Cluster Discovery

Run these commands to understand the cluster setup:

  • Check cluster details: gcloud container clusters describe {cluster_name} --location {location} --project {project}

  • Check for Autopilot vs Standard: Look for the following block in the describe output:

    autopilot:
      enabled: true
    
  • Check release channel: Look for releaseChannel.

Workload Discovery

If a specific application is targeted, discover its configuration:

  • Get deployment/statefulset details: kubectl get deployment {app_name} -n {namespace} -o yaml
  • Check for dedicated namespace and labels: kubectl get namespace {namespace} -o yaml (Look for Pod Security Standards labels).
  • Check for dedicated service account usage: kubectl get pods -n {namespace} -o custom-columns="NAME:.metadata.name,SERVICE_ACCOUNT:.spec.serviceAccountName"
  • Check for resource requests and limits.
  • Check for liveness, readiness, and startup probes.
  • Check for HPA: kubectl get hpa -n {namespace}
  • Check for PDB: kubectl get pdb -n {namespace}
  • Check for NetworkPolicies: kubectl get networkpolicy -n {namespace}

2. Production Readiness Assessment

Before implementation, you MUST run the skills for each relevant specialized area listed below and incorporate its guidance into your assessment and plan. Failure to do so will result in a non-compliant production configuration.

A. App Onboarding (Pre-Kubernetes)

If the application is not yet running on GKE, you MUST run the gke-app-onboarding skill for planning containerization, image building, and basic deployment.

B. Scalability & Resource Management

Ensure workloads have appropriate resources and autoscaling.

  • Action: You MUST run the gke-workload-scaling skill for configuring HPA, VPA, and resource limits.

C. Observability

Ensure adequate logging and monitoring are in place.

  • Action: You MUST run the gke-observability skill for setting up Cloud Logging, Monitoring, and Managed Prometheus.

D. Reliability

Ensure high availability and graceful degradation.

  • Action: You MUST run the gke-reliability skill for configuring regional clusters, PDBs, and health probes.

E. Security

Harden the cluster and workloads.

  • Action: You MUST run the gke-platform-security and gke-workload-security skills for Workload Identity, Network Policies, and Shielded Nodes.
  • Namespace Isolation: Ensure workloads run in dedicated namespaces with Pod Security Standards (PSS) enforced via labels.
  • Least Privilege: Ensure workloads use dedicated ServiceAccounts instead of the default ServiceAccount.

F. Backup & Disaster Recovery

Ensure stateful data is protected.

  • Action: You MUST run the gke-backup-dr skill for configuring Backup for GKE and restore procedures.

G. Edge Security & Ingress

Secure external access.

  • Action: You MUST run the gke-service-networking skill for Gateway API, Ingress, and Cloud Armor.

H. Cost Optimization

Ensure efficient use of resources.

  • Action: You MUST run the gke-cost-optimization skill for strategies on rightsizing, quotas, and Spot VMs.

I. Upgrades & Maintenance Posture

Ensure a safe, predictable upgrade posture.

  • Action: You MUST run the gke-upgrades skill for release channel selection, maintenance windows/exclusions, and node pool upgrade strategy.

J. Golden Path Defaults Audit

Ensure the cluster configuration matches recommended defaults.

  • Action: You MUST run the gke-golden-path skill to compare the cluster against golden path defaults and report deviations with severity and remediation.

3. Production Readiness Scoring

After the assessment, provide a summary report with a RAG (Red, Amber, Green) status for each area and an overall readiness score. This helps prioritize remediation efforts.

Apply this rubric deterministically so repeated assessments of the same environment produce the same result:

  1. Per-domain criteria: For each assessed domain (A-J), list the concrete checks performed (from the domain skill's guidance) and classify each check as pass, fail-critical (production-blocking, e.g., no resource requests, no backups for stateful data, public control plane in a locked down environment), or fail-minor (improvement, e.g., missing VPA recommendations, no Spot usage for batch).
  2. RAG mapping (per domain):
    • Red = one or more fail-critical checks.
    • Amber = no fail-critical, but one or more fail-minor checks.
    • Green = all checks pass.
  3. Domain score: Green = 100, Amber = 50, Red = 0.
  4. Weighted overall score: weight Security, Reliability, and Backup/DR at 2x; all other assessed domains at 1x. Overall score = sum(domain score x weight) / sum(weights), rounded to the nearest integer. Exclude domains that are not applicable (e.g., Backup/DR for fully stateless workloads) from both sums and note the exclusion.
  5. Readiness verdict: >= 90 with no Red domains = "Production ready"; 70-89 with no Red domains = "Ready with follow-ups"; anything else = "Not production ready".

In the report, show the per-domain check lists, RAG status, weights, and the computed overall score.

Adaptability Guidelines

  • Single App: Focus on Health Probes, HPA, Resource Limits, PDB, and Workload Identity for that specific app.
  • Cluster Wide: Focus on Cluster Autoscaler, Multi-zonal setup, Release Channels, Maintenance Windows, and default Network Policies.
  • Proactive Execution: Proactively execute relevant skills (e.g., observability, security, scaling, reliability) to assess and propose improvements, seeking user confirmation before applying state-changing implementations.

Files

1
7.2 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from google/skills8

agent-platform-alert-configuration

Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st

Needs review 0
agent-platform-deploy

Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model

Scan passed 0
agent-platform-endpoint-management

Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run

Scan passed 0
agent-platform-eval-flywheel

Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be

Scan passed 0
agent-platform-inference

Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c

Scan passed 0
agent-platform-migrate-from-ai-studio

Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry

Scan passed 0
agent-platform-model-registry

Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.

Scan passed 0
agent-platform-prompt-management

Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.

Scan passed 0

Related security skillsscan passed

security-review

AI-powered codebase security scanner that reasons about code like a security researcher — tracing data flows, understanding component interactions, and catching vulnerabilities that pattern-matching tools miss. Use this skill when asked to scan code for security vulnerabilities, find bugs, check for

Scan passed 1
security-threat-model

Repository-grounded threat modeling that enumerates trust boundaries, assets, attacker capabilities, abuse paths, and mitigations, and writes a concise Markdown threat model. Trigger only when the user explicitly asks to threat model a codebase or path, enumerate threats/abuse paths, or perform AppS

Scan passed 1
dotnet-patterns

Idiomatic C# and .NET patterns, conventions, dependency injection, async/await, and best practices for building robust, maintainable .NET applications. Use when writing or reviewing C# / .NET code — DI, async, or general conventions.

Scan passed 0
cso

Security audit: supported static findings; qualified profiles add reproduction and repair candidates. (gstack)

Scan passed 0
claude-security

Claude Security: scan the codebase (the whole repository or a scoped part of it), scan changes (this branch's or a pull request's diff, or one commit), or suggest patches (findings turned into targeted patch files, each verified by a panel of agents, that you apply when you choose). Use when the use

Scan passed 0
auth

Implement JWT/cookie authentication and authorization in tRPC using createContext for user extraction, t.middleware with opts.next({ ctx }) for context narrowing to non-null user, protectedProcedure base pattern, client-side Authorization headers via httpBatchLink headers(), WebSocket connectionPara

Scan passed 0