What problem does it solve?
Designing system availability from SLA targets is hard to keep consistent across layers, so this Skill helps produce a complete availability design that aligns SLA/uptime targets with RTO/RPO and practical redundancy and failure-recovery mechanisms.
Core Features & Use Cases
- SLA and availability target derivation: Defines uptime goals (SLA and uptime rate), and clarifies how planned maintenance is treated in the calculation.
- SPOF analysis and redundancy planning: Creates a structured single points of failure inventory and maps each component to a redundancy strategy or an explicit risk rationale.
- End-to-end resilience design: Specifies multi-AZ/multi-layer redundancy for compute, database, cache, storage, network, and DNS, including failover, auto scaling, health checks, planned stop/maintenance, and verification steps.
Use cases include producing or updating “availability design documents,” defining redundancy/failover/auto scaling, and generating design artifacts when asked to “design for zero-downtime deployments” or “eliminate SPOFs.”
Quick Start
Generate an availability design document that defines SLA and uptime targets, performs SPOF analysis, and specifies multi-AZ redundancy, failover, auto scaling, health checks, planned maintenance, and verification for the compute, DB, cache, storage, network, and DNS layers.