runbook

Create operational runbooks with health checks, escalation paths, and recovery steps.

19|3|Updated Feb 28, 2026
One-click install
npx skills add https://github.com/qa-aman/claude-skills --skill runbook-qa-aman
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: runbook
Source: https://github.com/qa-aman/claude-skills/tree/main/skills/by-role/devops/runbook
Command: npx skills add https://github.com/qa-aman/claude-skills --skill runbook-qa-aman

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Operational runbooks ensure consistent, executable guidance for on-call engineers during incidents, reducing guesswork and human error.

Core Features & Use Cases

  • Clear ownership, health checks, SLOs, and escalation paths to standardize incident response.
  • Step-by-step recovery procedures, including restart, scale, and rollback actions, plus explicit command examples.
  • Documentation templates that can be reused across services and teams to accelerate onboarding.

Quick Start

Provide a complete, executable runbook for a sample service so an on-call engineer can follow it during a simulated outage.

Frequently Asked Questions about runbook

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I create an operational runbook for incident response?

To create an operational runbook, define clear ownership, health checks, SLOs, and escalation paths, then document step-by-step recovery procedures with explicit command examples for on-call engineers to follow during incidents.

What should be included in a production service recovery runbook?

A production service recovery runbook should include executable procedures for restart, scale, and rollback actions, health checks, metrics references, clearly defined steps, and escalation contacts to guide on-call engineers.

Can I use runbook templates for routine maintenance and post-incident reviews?

Yes, runbook templates can be reused across services and teams for routine maintenance, post-incident reviews, and incident response documentation, providing standardized guidance that accelerates onboarding and reduces human error.

What is the best way to standardize on-call escalation paths and procedures?

The best way to standardize on-call escalation paths is to document executable procedures with health checks, metrics references, defined steps, and escalation contacts in reusable templates across all production services.

Does this approach support rollback and scale actions for simulated outages?

Yes, this approach supports rollback and scale actions by providing a complete, executable runbook with explicit command examples that an on-call engineer can follow during a simulated outage or actual incident.