skills/ HoangNguyen0403/agent-skills-standard

system-design-resilience-ops

Make a design survivable and operable: eliminate single points of failure, pick failover topology, set RPO/RTO, define readiness signals and observability, choose a deployment and rollback strategy. Use when reviewing availability, planning DR, or deciding rollout mechanics.

0
Installs
—
Rating
—
Success rate
3
Files scanned
Scan passedfrontend
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

3 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 e957576482819d70… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Resilience and Operations

Priority: P1 (HIGH)

A design is not done until its failure and its rollout are designed.

SPOF Elimination

  • Walk every component and ask what happens when exactly one instance dies, then when the whole zone dies.
  • Any component with one instance, one writer, or one shared config plane is a single point of failure. Name it or remove it.
  • Redundancy only helps when failure modes are independent: shared credentials, shared config, and a shared control plane cancel the benefit.
  • Blast radius: state which users or flows are affected per component failure, and cap it with cells, bulkheads, or per-tenant quotas.

Failover and Recovery

TopologyRecovery timeCostFits
Single region, multi-AZMinutes, automaticLowMost products
Active-passive across regionsMinutes to hours, drill-dependentMediumRegulated or high-value flows
Active-active across regionsSecondsHighGlobal low-latency, conflict-tolerant data
  • Set RPO (tolerable data loss) and RTO (tolerable downtime) as numbers before choosing a topology; the numbers pick the topology, not the reverse.
  • Untested failover is a hypothesis. Schedule a drill and record the measured RTO against the target.
  • Backups need a restore test. A backup that has never been restored is not a backup.

Observability

  • Instrument the four signals per service: traffic, error rate, latency percentiles, saturation.
  • Alert on user-visible symptoms and on error-budget burn rate, not on raw CPU.
  • Propagate a trace and correlation id across every hop, including queue messages.
  • Every alert needs an owner, a runbook link, and a defined next action; an alert nobody acts on is noise.

Rollout

StrategyBlast radiusRollbackCost
RollingGrows during the rollRoll forward or back, slowLow
Blue-greenFull switch at cutoverInstant switch backDouble capacity
CanarySmall cohort firstStop and drain the cohortNeeds routing plus metrics
Feature flagPer user or tenantInstant, no redeployFlag lifecycle debt
  • Schema and code deploy separately: expand, migrate, contract. Never ship a migration that only the new code can read.
  • Define the rollback trigger as a metric threshold and a time box before the deploy starts.

Anti-Patterns

  • No untested failover: no DR claim without a drill date and a measured RTO.
  • No unbounded retry: retries need budget, backoff with jitter, and a stop condition, or they amplify an outage.
  • No liveness probe on dependencies: a downstream outage must not restart the fleet.
  • No deploy without rollback: irreversible releases are outages waiting for a bad build.
  • No autoscaling without a floor and ceiling: unbounded scaling turns a bug into a bill.

References

Files

3
11.7 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from HoangNguyen0403/agent-skills-standard8

android-agp-upgrade

Upgrade an Android project to Android Gradle Plugin (AGP) 9. Use when migrating to AGP 9, updating Gradle build files, migrating to built-in Kotlin, or adopting the new AGP DSL.

Scan passed 0
android-architecture

Apply Clean Architecture layering, modularization, and Unidirectional Data Flow in Android projects. Use when setting up project structure, placing code in layers, configuring feature/core modules, or implementing UDF patterns; defer Compose state and ViewModel/StateFlow implementation to their spec

Scan passed 0
android-background-work

Implement WorkManager and background processing correctly on Android. Use when creating Worker classes, scheduling tasks, choosing between WorkManager and Foreground Services, or setting up Hilt in workers; defer FCM and notification delivery to android-notifications.

Scan passed 0
android-compose

Build high-performance declarative UI with Jetpack Compose. Use when writing Composable functions, optimizing recomposition, hoisting state, or working with LazyColumn and side effects; defer deep-link and navigation routing to android-navigation.

Scan passed 0
android-compose-migration

Migrate an Android XML View to Jetpack Compose following a structured 10-step workflow. Use when converting XML layouts to Compose, setting up Compose in an existing View-based project, or incrementally adopting Compose.

Scan passed 0
android-concurrency

Write correct coroutine scopes, lifecycle collection, and dispatcher injection in Android production code. Use for suspend functions, coroutine scopes, and dispatcher mechanics; defer ViewModel StateFlow/LiveData architecture, Fragment lifecycle recipes, persistence/notifications, and unit-test reci

Scan passed 0
android-deployment

Configure release signing, R8 obfuscation, and App Bundle publishing for Android. Use when setting up signing configs, enabling minification, adding ProGuard keep rules, or preparing for Play Store submission.

Scan passed 0
android-design-system

Enforce Material Design 3 theming and design token usage in Jetpack Compose. Use when implementing M3 components, color schemes, typography, or design tokens.

Scan passed 0

Related frontend skillsscan passed