system-design-resilience-ops
Make a design survivable and operable: eliminate single points of failure, pick failover topology, set RPO/RTO, define readiness signals and observability, choose a deployment and rollback strategy. Use when reviewing availability, planning DR, or deciding rollout mechanics.
- 0
- Installs
- —
- Rating
- —
- Success rate
- 3
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 e957576482819d70… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Resilience and Operations
Priority: P1 (HIGH)
A design is not done until its failure and its rollout are designed.
SPOF Elimination
- Walk every component and ask what happens when exactly one instance dies, then when the whole zone dies.
- Any component with one instance, one writer, or one shared config plane is a single point of failure. Name it or remove it.
- Redundancy only helps when failure modes are independent: shared credentials, shared config, and a shared control plane cancel the benefit.
- Blast radius: state which users or flows are affected per component failure, and cap it with cells, bulkheads, or per-tenant quotas.
Failover and Recovery
| Topology | Recovery time | Cost | Fits |
|---|---|---|---|
| Single region, multi-AZ | Minutes, automatic | Low | Most products |
| Active-passive across regions | Minutes to hours, drill-dependent | Medium | Regulated or high-value flows |
| Active-active across regions | Seconds | High | Global low-latency, conflict-tolerant data |
- Set RPO (tolerable data loss) and RTO (tolerable downtime) as numbers before choosing a topology; the numbers pick the topology, not the reverse.
- Untested failover is a hypothesis. Schedule a drill and record the measured RTO against the target.
- Backups need a restore test. A backup that has never been restored is not a backup.
Observability
- Instrument the four signals per service: traffic, error rate, latency percentiles, saturation.
- Alert on user-visible symptoms and on error-budget burn rate, not on raw CPU.
- Propagate a trace and correlation id across every hop, including queue messages.
- Every alert needs an owner, a runbook link, and a defined next action; an alert nobody acts on is noise.
Rollout
| Strategy | Blast radius | Rollback | Cost |
|---|---|---|---|
| Rolling | Grows during the roll | Roll forward or back, slow | Low |
| Blue-green | Full switch at cutover | Instant switch back | Double capacity |
| Canary | Small cohort first | Stop and drain the cohort | Needs routing plus metrics |
| Feature flag | Per user or tenant | Instant, no redeploy | Flag lifecycle debt |
- Schema and code deploy separately: expand, migrate, contract. Never ship a migration that only the new code can read.
- Define the rollback trigger as a metric threshold and a time box before the deploy starts.
Anti-Patterns
- No untested failover: no DR claim without a drill date and a measured RTO.
- No unbounded retry: retries need budget, backoff with jitter, and a stop condition, or they amplify an outage.
- No liveness probe on dependencies: a downstream outage must not restart the fleet.
- No deploy without rollback: irreversible releases are outages waiting for a bad build.
- No autoscaling without a floor and ceiling: unbounded scaling turns a bug into a bill.
References
- Reliability Operations - failure drills, health check design, DR runbook shape, scaling policy notes
Files
3- SKILL.md
0d9f714d903.5 KB - evals/evals.json
f6173142574.8 KB - references/reliability-operations.md
3a9eac9d6b3.4 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from HoangNguyen0403/agent-skills-standard8
Upgrade an Android project to Android Gradle Plugin (AGP) 9. Use when migrating to AGP 9, updating Gradle build files, migrating to built-in Kotlin, or adopting the new AGP DSL.
Apply Clean Architecture layering, modularization, and Unidirectional Data Flow in Android projects. Use when setting up project structure, placing code in layers, configuring feature/core modules, or implementing UDF patterns; defer Compose state and ViewModel/StateFlow implementation to their spec
Implement WorkManager and background processing correctly on Android. Use when creating Worker classes, scheduling tasks, choosing between WorkManager and Foreground Services, or setting up Hilt in workers; defer FCM and notification delivery to android-notifications.
Build high-performance declarative UI with Jetpack Compose. Use when writing Composable functions, optimizing recomposition, hoisting state, or working with LazyColumn and side effects; defer deep-link and navigation routing to android-navigation.
Migrate an Android XML View to Jetpack Compose following a structured 10-step workflow. Use when converting XML layouts to Compose, setting up Compose in an existing View-based project, or incrementally adopting Compose.
Write correct coroutine scopes, lifecycle collection, and dispatcher injection in Android production code. Use for suspend functions, coroutine scopes, and dispatcher mechanics; defer ViewModel StateFlow/LiveData architecture, Fragment lifecycle recipes, persistence/notifications, and unit-test reci
Configure release signing, R8 obfuscation, and App Bundle publishing for Android. Use when setting up signing configs, enabling minification, adding ProGuard keep rules, or preparing for Play Store submission.
Enforce Material Design 3 theming and design token usage in Jetpack Compose. Use when implementing M3 components, color schemes, typography, or design tokens.
Related frontend skillsscan passed
Compose Multiplatform and Jetpack Compose patterns for KMP projects — state management, navigation, theming, performance, and platform-specific UI. Use when building Compose or Jetpack Compose UI, state, navigation, or theming in a KMP project.
Build UIs with @nuxt/ui v4 — 125+ accessible Vue components with Tailwind CSS theming. Use when creating interfaces, customizing themes to match a brand, building forms, or composing layouts like dashboards, docs sites, and chat interfaces.
Import cookies from your real Chromium browser into the headless browse session. (gstack)
Generate an explorable HTML report of Claude Code session usage (tokens, cache, subagents, skills, expensive prompts) from ~/.claude/projects transcripts.
Renders a component you choose under every scenario that can reach it on a temporary page and stress tests it.
Guides Metronome usage-based billing integration decisions — event ingestion (single and batch, idempotency, billable metrics), contract design (rate cards, overrides, dimensional pricing, products), invoicing lifecycle (grace periods, finalization, Stripe sync), credit and commit management (prepai