role-devops:incident-response

Create runbooks, postmortem templates, and status page updates for incident response.

14|3|Updated Feb 22, 2026
One-click install
npx skills add https://github.com/rnavarych/alpha-engineer --skill role-devops-incident-response
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: role-devops:incident-response
Source: https://github.com/rnavarych/alpha-engineer/tree/main/plugins/roles/role-devops/skills/incident-response
Command: npx skills add https://github.com/rnavarych/alpha-engineer --skill role-devops-incident-response

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill provides comprehensive guidance and templates for establishing robust incident response processes, ensuring faster recovery and improved system resilience.

Core Features & Use Cases

  • Runbook Creation: Develop detailed, step-by-step guides for alert remediation.
  • Postmortem Analysis: Conduct blameless postmortems to identify root causes and prevent recurrence.
  • On-Call & Escalation: Design effective on-call rotations and escalation policies.
  • Chaos Engineering: Implement proactive resilience testing through chaos experiments and game days.
  • Status Pages: Manage public status pages for transparent customer communication.
  • Use Case: A new S1 incident has occurred. Use this Skill to generate a blameless postmortem template, define severity levels, and draft an update for the public status page.

Quick Start

Use the incident-response skill to create a blameless postmortem template.

Frequently Asked Questions about role-devops:incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I write a blameless postmortem after a major incident?

A blameless postmortem identifies root causes and prevents recurrence without assigning individual blame. This Skill provides structured postmortem templates to systematically document incidents, define severity levels, and extract actionable remediation items.

How do I set up an on-call rotation and escalation policy for incident response?

Effective on-call rotations require clear schedules and escalation paths to route critical alerts correctly. This Skill helps design on-call rotations and escalation policies, ensuring the right responders are notified during system failures.

What's the best way to create runbooks for alert remediation?

Runbooks provide step-by-step guides for alert remediation to ensure consistent and rapid recovery. This Skill facilitates runbook creation, helping you develop detailed, repeatable processes for handling operational alerts.

How does chaos engineering improve system resilience?

Chaos engineering proactively tests system resilience by injecting failures before they cause real incidents. This Skill supports implementing chaos experiments and organizing game day exercises to validate your infrastructure's reliability.

How do I manage a public status page during an S1 incident?

Managing a public status page ensures transparent customer communication during critical events like an S1 incident. This Skill helps you draft status updates and manage public communication channels while the incident response is active.

Can I use this for establishing incident response processes from scratch?

Yes, establishing incident response processes requires clear documentation and repeatable workflows from the start. This Skill provides comprehensive expertise and templates for severity definitions, postmortems, and runbooks to build your process foundation.