What problem does it solve?
Organizations struggle to standardize resilience policies across teams and portfolios, often creating inconsistent availability targets, mismatched DR approaches, and ad-hoc schedules for resilience testing activities.
Core Features & Use Cases
- Tiered Policy Model: Classifies services by business criticality and maps each tier to availability SLO, RTO/RPO targets, and a DR approach (e.g., ACTIVE_ACTIVE, WARM_STANDBY, BACKUP_AND_RESTORE) validated against the Resilience Hub API enums.
- Operational Cadence Guidance: Recommends how often to run resilience activities, from continuous ARC zonal autoshift practice runs to weekly findings reviews, monthly FIS experiments, and quarterly cross-service GameDays.
- Security Standards: Embeds least-privilege IAM roles, short-lived credentials, SSE-KMS encryption, and FIS production governance into program-level standards.
- Use Case: A platform lead asks how to structure resilience policies for 50 workloads; the Skill recommends a three-tier model with payments at 99.99 availability and ACTIVE_ACTIVE DR, and dev/test at 99.9 with BACKUP_AND_RESTORE.
Quick Start
Ask the agent to design a tiered resilience policy model with availability, RTO/RPO targets, and a testing cadence for your organization's workloads.