disaster-recovery-exercise-design

Design and run disaster recovery exercises with tabletop and live scenarios.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/ohsonerdy/openclaw-frontier-stack --skill disaster-recovery-exercise-design
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: disaster-recovery-exercise-design
Source: https://github.com/ohsonerdy/openclaw-frontier-stack/tree/main/skills/disaster-recovery-exercise-design
Command: npx skills add https://github.com/ohsonerdy/openclaw-frontier-stack --skill disaster-recovery-exercise-design

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps teams design disaster recovery (DR) exercises that actually validate whether their recovery plan works under realistic failure conditions, including measurable RTO and RPO outcomes.

Core Features & Use Cases

  • DR exercise planning (tabletop to live): Choose the right drill format (tabletop, partial live, or full live) based on risk, maturity, and blast radius.
  • Failure-mode and blast-radius selection: Select which high-value failure modes to rehearse and frame expected worst-case impact with clear abort criteria and customer notification plans.
  • Success criteria and evidence discipline: Define measurable RTO/RPO targets, validate runbook reality, verify paging/on-call response, and enforce action-item follow-through with owners and deadlines.

Quick Start

Ask for a DR drill plan that targets a specific failure mode (for example, primary database loss) with defined RTO/RPO success criteria, tabletop-to-live escalation, participant roles, data collection steps, and an action-item tracking approach.

Frequently Asked Questions about disaster-recovery-exercise-design

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design a disaster recovery exercise that actually validates RTO and RPO?▼

Disaster recovery exercise design validates RTO and RPO by selecting specific failure modes, defining measurable success criteria, and rehearsing region failover or backup-restore scenarios. It structures drills through tabletop to live escalation with explicit data collection and action-item tracking.

What is the best way to run a tabletop disaster recovery drill for region failover?▼

Tabletop disaster recovery drills for region failover start with scenario framing and blast radius selection. You define worst-case impact, abort criteria, and customer notification plans, then validate runbook reality and on-call paging response with assigned participant roles.

How do I validate on-call paging and runbook reality during incident response rehearsal?▼

Incident response rehearsal validates on-call paging and runbook reality by simulating high-value failure modes and measuring actual response times against RTO/RPO targets. It enforces follow-through via owned action items with deadlines to ensure operational governance.

When should I choose a partial live DR drill over a full live chaos engineering scenario?▼

Choose a partial live DR drill over full live chaos engineering based on risk, maturity, and blast radius. Disaster recovery exercise design scales the format from tabletop to full live to safely test backup-restore validation while maintaining explicit abort criteria.

Can I use disaster recovery exercise planning to test team readiness for primary database loss?▼

Disaster recovery exercise planning tests team readiness for primary database loss by framing the scenario, selecting the failure mode, and defining RTO/RPO measurement criteria. It structures participant timing and data collection to validate operational governance.

What limitations should I anticipate when scheduling live DR drills for backup-restore validation?▼

Live DR drills for backup-restore validation carry blast radius risks requiring clear abort criteria and customer notification plans. Scheduling must account for participant availability, timing constraints, and the maturity of existing runbooks before executing full live failover.