What problem does it solve?
Teams struggle to connect resilience policy definition, fault-injection testing, and operational recovery controls into one coherent program, often marking findings resolved without proving fixes work. This Skill guides the end-to-end Define → Test → Operate workflow so resilience findings are validated with real experiments before closure.
Core Features & Use Cases
- Integrated Lifecycle Orchestration: Sequences Resilience Hub v2 (NGRH) policy creation and failure-mode assessment, FIS experiment validation, and ARC routing controls and zonal autoshift into a gated workflow.
- Hallucination-Resistant API Reference: Provides canonical AWS CLI operations for resiliencehubv2, fis, route53-recovery-control-config, and arc-zonal-shift, with a table of common wrong API names mapped to correct ones.
- Validation-First Finding Resolution: Enforces that NGRH findings are marked RESOLVED only after an FIS experiment confirms recovery within RTO/RPO targets, preventing paper compliance.
- Use Case: A platform team wants to know if they are "done" after resolving Resilience Hub findings. The Skill walks them through validating each fix with a fault-injection experiment, then setting up ARC routing controls and zonal autoshift for ongoing operational resilience.
Quick Start
Ask the agent to guide you through the full AWS resilience lifecycle for your service, from creating a Resilience Hub policy through FIS experiment validation to ARC operational controls.