google-cloud-waf-performance-optimization
Generates performance-focused guidance for Google Cloud workloads based on the design principles and recommendations in the Performance Optimization pillar of the Google Cloud Well-Architected Framework (WAF). Use this skill to evaluate a workload, identify performance requirements, and provide acti
- 0
- Installs
- —
- Rating
- —
- Success rate
- 1
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 d888c5453a88cf37… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Google Cloud Well-Architected Framework skill for the Performance Optimization pillar
Overview
The Performance Optimization pillar of the Google Cloud Well-Architected Framework provides principles and recommendations to help you design, build, and operate high-performing workloads. It focuses on efficiently allocating resources, leveraging modular architectures, and using data-driven insights to continuously monitor and improve performance as your business needs evolve.
Core principles
The recommendations in the performance optimization pillar of the Well-Architected Framework are aligned with the following core principles:
-
Plan resource allocation: Carefully select and configure the compute, storage, and networking resources that best match the specific requirements of your workload. Grounding document: https://docs.cloud.google.com/architecture/framework/performance-optimization/plan-resource-allocation.md.txt
-
Take advantage of elasticity: Utilize automated scaling and serverless technologies to dynamically adjust resource capacity in response to real-time demand fluctuations. Grounding document: https://docs.cloud.google.com/architecture/framework/performance-optimization/elasticity.md.txt
-
Promote modular design: Architect systems using independent, loosely coupled components to enhance scalability and allow individual parts to be optimized without affecting the entire system. Grounding document: https://docs.cloud.google.com/architecture/framework/performance-optimization/promote-modular-design.md.txt
-
Continuously monitor and improve performance: Implement robust observability to identify bottlenecks and use performance data to drive iterative enhancements throughout the software development lifecycle. Grounding document: https://docs.cloud.google.com/architecture/framework/performance-optimization/continuously-monitor-and-improve-performance.md.txt
Relevant Google Cloud products
The following are examples of Google Cloud products and features that are relevant to performance optimization:
-
Compute and scaling
- Compute Engine (MIGs): Managed instance groups that support autoscaling and load balancing for VM-based workloads.
- Google Kubernetes Engine (GKE): Provides container orchestration with horizontal and vertical pod autoscaling.
- Cloud Run: A fully managed serverless platform that automatically scales containers to zero or up based on traffic.
-
Data and caching
- Cloud CDN: Low-latency content delivery network to cache static and dynamic content closer to end-users.
- Memorystore: Managed in-memory data store for Valkey and Redis to provide sub-millisecond data access.
- Bigtable: NoSQL database service for analytical and operational workloads requiring low latency and high throughput.
- Spanner: RDBMS that provides global consistency, high availability, and horizontal scaling for mission-critical transactional applications.
-
Performance analysis and monitoring
- Cloud Trace: Distributed tracing system that helps identify latency bottlenecks.
- Cloud Profiler: Continuous CPU and memory profiling to identify resource-heavy application code.
- Cloud Monitoring: Provides dashboards and alerts based on performance KPIs like latency and throughput.
Workload assessment questions
Ask appropriate questions to understand the performance-related requirements and constraints of the workload and the user's organization. Choose questions from the following list:
-
Plan resource allocation
- When initially provisioning compute resources for a new application, which approach do you use to determine the required capacity for expected peak loads?
- Which caching strategies (browser, in-memory, CDN, database) do you utilize to improve performance and responsiveness?
- How do you optimize the performance of your data storage solutions (e.g., SSD vs HDD, storage classes) for your applications?
-
Promote modular design
- Which architectural patterns (microservices, asynchronous messaging, stateless servers) do you employ to enhance performance and resilience?
- How do you design your application to minimize the impact of failures in one part of the system on other parts?
-
Continuously monitor and improve performance
- How frequently do you review and analyze the performance of your production applications and infrastructure?
- Which tools or techniques (APM, distributed tracing, load testing) do you use to proactively identify and diagnose performance bottlenecks?
- How do you incorporate performance considerations into your software development lifecycle (SDLC)?
-
Take advantage of elasticity
- Which methods do you use to manage and optimize the cost of your cloud resources while maintaining performance?
- How do you typically handle sudden spikes in traffic or workload on your applications?
Validation checklist
Use the following checklist to evaluate the architecture's alignment with performance optimization recommendations:
-
Resource allocation
- Initial provisioning is based on load testing or historical data rather than general estimates.
- Caching is implemented at multiple layers (CDN, in-memory, or browser) to offload backend systems.
- Storage types (SSD/HDD) and classes are selected based on the specific I/O requirements of the workload.
-
Modular design
- The architecture uses microservices or decoupled components to allow independent scaling.
- Circuit breakers or bulkheads are implemented to isolate failures and prevent performance degradation across the system.
-
Monitoring and continuous improvement
- Automated dashboards and alerts are configured for key performance indicators (KPIs).
- Distributed tracing and profiling tools are used to identify code-level bottlenecks.
- Performance testing (unit and integration) is integrated into the software development lifecycle.
-
Elasticity
- Auto-scaling rules are configured and validated to handle variable demand.
- The architecture leverages serverless or managed services to dynamically match capacity to load.
- Resource utilization is reviewed regularly to eliminate idle overhead and balance cost with performance.
Files
1- SKILL.md
02a2d0b0227.2 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from google/skills8
Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st
Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model
Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run
Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be
Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c
Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry
Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.
Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.
Related tooling skillsscan passed
Design and optimize AI agent action spaces, tool definitions, and observation formatting for higher completion rates. Use when defining or revising an agent's tool set, action space, or observation format.
Web performance regression detection. (gstack)
This skill should be used when the user asks to "demonstrate skills", "show skill format", "create a skill template", or discusses skill development patterns. Provides a reference template for creating Claude Code plugin skills.
Helps you build and check a color system for your project. It generates palettes, names semantic tokens, converts between formats and measures contrast.
Creates a new Angular app using the Angular CLI. This skill should be used whenever a user wants to create a new Angular application and contains important guidelines for how to effectively create a modern Angular application.
Audit, diagnose, or optimize website loading and interaction performance, Core Web Vitals, and Lighthouse performance scores.