on-call-runbook

Generate structured on-call runbooks for service alerts with triage steps and remediation.

1|Updated Mar 6, 2026
One-click install
npx skills add https://github.com/chavangorakh1999/sde-skills --skill on-call-runbook
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: on-call-runbook
Source: https://github.com/chavangorakh1999/sde-skills/tree/main/sde-execution/skills/on-call-runbook
Command: npx skills add https://github.com/chavangorakh1999/sde-skills --skill on-call-runbook

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill helps create comprehensive on-call runbooks that guide engineers through diagnosing and resolving alerts for a service, reducing downtime and improving incident response.

Core Features & Use Cases

  • Structured Runbook Generation: Provides a template for documenting alerts, probable causes, triage steps, and fixes.
  • Incident Response Guidance: Ensures consistent and efficient handling of production incidents.
  • Use Case: When a new alert fires for a critical service, use this Skill to quickly generate a runbook that details how to respond, what to check, and when to escalate.

Quick Start

Use the on-call runbook skill to generate a runbook for the 'HighErrorRate' alert on the 'user-auth-service'.

Frequently Asked Questions about on-call-runbook

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I create an on-call runbook for production service alerts?

To create an on-call runbook for production service alerts, provide detailed service context and alert specifics to generate structured documentation covering alert conditions, triage steps, probable causes, and remediation procedures.

What should be included in an incident response runbook for DevOps troubleshooting?

An incident response runbook for DevOps troubleshooting should include alert conditions, triage steps, probable causes, and remediation procedures to ensure consistent handling of production incidents and reduce service downtime.

Can I generate troubleshooting steps for any type of service alert?

Yes, you can generate troubleshooting steps for any service alert by supplying detailed input on the specific service context and alert specifics, allowing the system to create actionable guidance for incident response.

What is the best way to document triage steps and remediation procedures for incident response?

The best way to document triage steps and remediation procedures is to use a structured runbook template that details alert conditions and probable causes, ensuring engineers have clear, actionable guidance during production incidents.

What input is required to generate an effective on-call runbook?

Generating an effective on-call runbook requires detailed input on service context and alert specifics, such as the target service name and the alert type, to accurately document triage and remediation procedures.

When do I need a structured runbook for alerting and incident response?

You need a structured runbook for alerting and incident response when a new alert fires for a critical service, requiring clear documentation on how to respond, what to check, and when to escalate to reduce downtime.