incident-playbook

Generate incident playbooks covering nine failure categories for production AI systems.

1|Updated Mar 26, 2026
One-click install
npx skills add https://github.com/selcukyucel/north-starr-genai --skill incident-playbook-selcukyucel
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-playbook
Source: https://github.com/selcukyucel/north-starr-genai/tree/main/skills/incident-playbook
Command: npx skills add https://github.com/selcukyucel/north-starr-genai --skill incident-playbook-selcukyucel

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Generate comprehensive, production-grade incident runbooks that standardize responses to nine AI failure modes and guide on-call teams through detection, containment, and recovery.

Core Features & Use Cases

  • Nine incident categories with structured runbooks covering detection, severity, response, root-cause analysis, resolution, prevention, and client communications.
  • On-demand generation for new automation pipelines and existing systems, enabling rapid recovery planning and SLA assurance.
  • Templates for stakeholder updates and escalation paths that align with incident severity and blast radius.

Quick Start

Provide the automation name to generate a complete incident playbook covering nine failure categories.

Frequently Asked Questions about incident-playbook

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I create an incident playbook for AI production systems?

To create an incident playbook for AI production systems, provide the automation name to generate structured runbooks covering nine failure categories. The output includes detection, severity, immediate response, root cause analysis, resolution, prevention, and client communication templates.

What should an on-call runbook cover for AI automation outages?

An on-call runbook for AI automation outages should cover detection, severity classification, immediate response, root cause analysis, resolution, prevention, and client communication templates. It standardizes responses to nine distinct AI failure modes to guide teams through containment and recovery.

How do incident playbooks handle escalation paths and stakeholder updates?

Incident playbooks handle escalation paths by providing templates for stakeholder updates that align with incident severity and blast radius. This ensures on-call teams can rapidly communicate recovery plans and SLA assurances during AI system disruptions.

Can I generate incident response runbooks for existing data pipelines?

Yes, you can generate incident response runbooks for existing data pipelines and new automation pipelines. The on-demand generation produces comprehensive runbooks covering nine AI failure categories, enabling rapid recovery planning and SLA assurance for production systems.

What are the nine failure categories addressed by AI incident runbooks?

The nine failure categories addressed by AI incident runbooks encompass production AI system failures involving external dependencies, data pipelines, monitoring, and on-call escalation. Each category includes structured runbooks for detection, containment, and recovery.

When do I need a structured incident playbook for my automation system?

You need a structured incident playbook for your automation system when it operates as a production AI system with external dependencies, data pipelines, monitoring, and on-call escalation. It enables rapid incident response and standardized recovery across nine failure modes.