skills/ google/skills

gke-backup-dr

Configures Backup for GKE: the BackupRestore cluster addon, BackupPlan and RestorePlan resources, restore workflows, and CMEK-encrypted backups. Use for backup policies, disaster recovery, or GKE cluster restores. Don't use for database backups.

0
Installs
—
Rating
—
Success rate
1
Files scanned
Scan passedsecurity
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

1 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 6e0d7e9f6ed67d8c… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

GKE Backup & Disaster Recovery

Protects stateful GKE workloads using Backup for GKE. Backup for GKE can capture both Kubernetes resource metadata (manifests, configurations, and secrets) and the underlying persistent volume (PV) data — but volume data and secrets are only captured when the backup plan explicitly enables them (see the flags below).

CLI Reference

# Enable the BackupRestore addon (Slow cluster-level update)
gcloud container clusters update {cluster_name} \
  --update-addons=BackupRestore=ENABLED --location={location} --quiet

# Create Backup Plan
gcloud beta container backup-restore backup-plans create {plan_name} \
  --project={project_id} --location={location} \
  --cluster=projects/{project_id}/locations/{location}/clusters/{cluster_name} \
  --all-namespaces \
  --include-volume-data --include-secrets \
  --backup-retain-days={days} --cron-schedule="{cron}" --quiet

# Trigger Manual Backup
gcloud beta container backup-restore backups create {backup_name} \
  --backup-plan={plan_name} --location={location} --quiet

# Create Restore Plan
gcloud beta container backup-restore restore-plans create {restore_plan_name} \
  --location={location} \
  --cluster=projects/{project_id}/locations/{location}/clusters/{target_cluster_name} \
  --backup-plan=projects/{project_id}/locations/{location}/backupPlans/{source_backup_plan_name} \
  --all-namespaces \
  --cluster-resource-conflict-policy=use-existing-version \
  --namespaced-resource-restore-mode=fail-on-conflict --quiet

# Execute Restore
gcloud beta container backup-restore restores create {restore_name} \
  --restore-plan={restore_plan_name} --location={location} \
  --backup=projects/{project_id}/locations/{location}/backupPlans/{source_backup_plan_name}/backups/{backup_name} \
  --quiet

# Verify Restore Status
gcloud beta container backup-restore restores describe {restore_name} \
  --restore-plan={restore_plan_name} --location={location}

[!WARNING] --include-volume-data and --include-secrets BOTH DEFAULT TO FALSE. If you omit them, the backup plan silently produces config-only backups with no persistent volume snapshots and no Secrets. Always pass both flags explicitly when the goal is full workload protection.

Notes:

  • The backup-restore command group requires the gcloud beta component (gcloud components install beta).
  • --cluster requires the full resource path projects/{project_id}/locations/{location}/clusters/{cluster_name} (or projects/{project_id}/zones/{zone}/clusters/{cluster_name} for zonal clusters), not a bare cluster name.
  • Restore plans require exactly one namespaced-resource scope flag: --all-namespaces, --selected-namespaces={ns1},{ns2}, --excluded-namespaces=..., --selected-applications=..., or --no-namespaces.

Restore Safety (CRITICAL)

A restore writes into a live cluster and, depending on the conflict policy, can overwrite or delete existing resources:

  • --cluster-resource-conflict-policy=use-existing-version keeps existing cluster-scoped resources (safe default); use-backup-version deletes the existing version first — deleting a CRD deletes all of its CRs.
  • --namespaced-resource-restore-mode=fail-on-conflict aborts on any conflict (safe default); merge-skip-on-conflict skips conflicting resources; merge-replace-on-conflict and merge-replace-volume-on-conflict overwrite existing resources or volumes; delete-and-restore deletes entire conflicting namespaces (and all resources in them) before restoring.

Rules:

  1. Validate the restore in a non-production target cluster first.
  2. Prefer the safe defaults (use-existing-version + fail-on-conflict) unless the user explicitly needs to revert live resources.
  3. Always obtain explicit user confirmation before executing a restore into a production cluster, and state which conflict policy is in effect and what it may overwrite or delete.

Best Practices

  1. CMEK Encryption: Encrypt backup plans using Customer-Managed Encryption Keys: --encryption-key=projects/{project_id}/locations/{location}/keyRings/{ring}/cryptoKeys/{key}.
  2. Scope: Prefer backing up specific namespaces rather than the entire cluster: --selected-namespaces={ns1},{ns2} (instead of --all-namespaces).
  3. Application Consistency: Recommend quiescing the database or pausing application writes (e.g. using pre-backup hooks or database-specific tools) prior to backups to ensure data integrity.
  4. CSI Volume Snapshots: Ensure that stateful backups utilize GKE's CSI (Container Storage Interface) driver for volume snapshots to capture persistent volume data.
  5. Service Terminology: Always explicitly refer to the service as Backup for GKE in your response. This distinguishes it from the broader (but complementary) Google Cloud Backup and Disaster Recovery (DR) Service, ## Golden Path Backup Defaults

The recommended production golden path configuration for Backup for GKE:

  • Addon: BackupRestore addon enabled (--update-addons=BackupRestore=ENABLED).
  • Volume Inclusion: --include-volume-data explicitly passed (enabled, since the service default is false).
  • Secret Inclusion: --include-secrets explicitly passed (enabled, since the service default is false).
  • Retention: Defined retention period (e.g. 30 days via --backup-retain-days=30).
  • Encryption: CMEK enabled (--encryption-key=...).

Recent Changes

  • Cross-project backup and restore (GA): Backup plans can store backups in a different project than the source cluster, and restore plans can target clusters in a third project. Enables centralized backup projects (with immutability/retention managed by a platform team) and cross-project environment seeding without granting access to the source project.
  • Pricing change (effective 2026-03-02): The backup management fee moved from pod-based to NAMESPACE-based pricing — charged per non-system namespace in the most recent successful backup of each plan (system namespaces like kube-system are excluded). Existing committed use discount (CUD) holders keep pod-based management pricing until their commitment ends; everyone else moves to the new model. See https://cloud.google.com/products/backup-for-gke/pricing-changes.
  • Smart Scheduling: RPO-driven backup scheduling as an alternative to fixed cron schedules — pass --target-rpo-minutes={minutes} instead of --cron-schedule when creating the backup plan (optionally with RPO exclusion windows via --exclusion-windows-file).
  • Hyperdisk support: Backup and restore of Hyperdisk ML and Hyperdisk Balanced High Availability volumes is supported on GKE clusters running 1.33.1-gke.1959000 and later (Hyperdisk throughput, extreme, and balanced types are also supported).

Troubleshooting & Common Pitfalls (CRITICAL)

[!IMPORTANT] Slow Operations: Enabling the BackupRestore addon (--update-addons=BackupRestore=ENABLED) triggers a slow Google Cloud control plane cluster update that takes several minutes. * Rule: Do not run a terminal loop waiting for the GKE Backup addon to become active. * Action: Provide the command to enable the addon, explain that the operation will proceed in the background, and immediately proceed to write the backup plan configs. Do not block.

Files

1
7.8 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from google/skills8

agent-platform-alert-configuration

Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st

Needs review 0
agent-platform-deploy

Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model

Scan passed 0
agent-platform-endpoint-management

Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run

Scan passed 0
agent-platform-eval-flywheel

Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be

Scan passed 0
agent-platform-inference

Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c

Scan passed 0
agent-platform-migrate-from-ai-studio

Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry

Scan passed 0
agent-platform-model-registry

Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.

Scan passed 0
agent-platform-prompt-management

Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.

Scan passed 0

Related security skillsscan passed

laravel-security

Laravel security best practices — authentication, authorization, Eloquent safety, CSRF, XSS prevention, API security, and secure deployment configurations. Use when reviewing Laravel auth, Eloquent safety, CSRF, XSS, API security, or deployment configuration.

Scan passed 0
cso

Security audit: supported static findings; qualified profiles add reproduction and repair candidates. (gstack)

Scan passed 0
claude-security

Claude Security: scan the codebase (the whole repository or a scoped part of it), scan changes (this branch's or a pull request's diff, or one commit), or suggest patches (findings turned into targeted patch files, each verified by a panel of agents, that you apply when you choose). Use when the use

Scan passed 0
client-setup

Create a vanilla tRPC client with createTRPCClient<AppRouter>(), configure link chain with httpBatchLink/httpLink, dynamic headers for auth, transformer on links (not client constructor). Infer types with inferRouterInputs and inferRouterOutputs. AbortController signal support. TRPCClientError typin

Scan passed 0
security-and-hardening

Hardens code against vulnerabilities. Use when auditing an input handler for vulnerabilities, when handling user input, authentication, data storage, or external integrations, or when checking a login flow is safe against the OWASP Top Ten. Use when building any feature that accepts untrusted data,

Scan passed 0
ponytail-audit

Quality audit of a whole repo: bugs, security holes, what breaks under real load, risky code without tests, slow paths, and what to delete, merge or split. Ranked, each finding explained in plain English. One-shot report, changes nothing. Use for "audit this codebase", "review the whole repo", "find

Scan passed 0