What problem does it solve?
This Skill helps teams design controlled operational drills that validate whether systems, alerts, runbooks, and responders can actually handle failures before a real incident happens. It prevents vague chaos exercises by forcing clear hypotheses, blast-radius limits, stop conditions, and measurable recovery goals.
Core Features & Use Cases
- Hypothesis-driven game day planning: Turns resilience testing into a falsifiable experiment with defined steady-state metrics, detection targets, and recovery objectives.
- Safety-first failure injection design: Guides teams through choosing realistic failure scenarios, minimizing blast radius, defining kill switches, and setting abort conditions.
- Operational readiness validation: Tests not just infrastructure behavior, but also alert quality, runbook accuracy, permissions, communications, and on-call response.
- Use Case: Before a major launch or peak-traffic event, use this Skill to plan a drill for database failover, dependency latency, or service degradation and capture action items from any gaps exposed.
Quick Start
Use the operational-game-day skill to design a game day plan for a critical service failover with hypotheses, stop conditions, roles, and post-drill findings tracking.