incident-response

Guide immediate mitigation actions for active production incidents.

1|Updated Jul 17, 2026
One-click install
npx skills add https://github.com/Arafly/sre-playbooks --skill incident-response-arafly
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-response
Source: https://github.com/Arafly/sre-playbooks/tree/main/incident-response
Command: npx skills add https://github.com/Arafly/sre-playbooks --skill incident-response-arafly

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps you respond correctly during an active incident when a service is broken, degraded, or at risk right now. It prioritizes fast mitigation and stabilization over premature root-cause analysis, reducing user impact before deeper diagnosis begins.

Core Features & Use Cases

  • Mitigate First: Recommends the fastest safe action such as rollback, failover, feature disablement, load shedding, or pausing risky jobs.
  • Stabilize and Scope: Confirms whether impact improved, then narrows the blast radius across users, regions, services, and workflows.
  • Preserve Evidence and Communicate: Captures a raw incident timeline, tracks actions and signals, and supports clear stakeholder updates throughout the event.
  • Use Case: A checkout outage, payment failure, production degradation, or data-risk event is detected and you need immediate response steps and a structured handoff to RCA.

Quick Start

Use the incident-response skill to help me mitigate this live production incident, stabilize impact, and produce a raw evidence-backed timeline.

Frequently Asked Questions about incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I mitigate a live production incident and reduce user impact?

To mitigate a live production incident, you must prioritize fast stabilization actions like rollback, failover, or feature disablement over premature root cause analysis, actively reducing user impact before deeper diagnosis begins.

What is the best way to handle an active service outage or degradation?

Handling an active service outage requires immediate incident scoping, evidence preservation, and mitigation validation to safely narrow the blast radius across affected users, regions, and workflows.

How do I preserve evidence during an active production outage?

To preserve evidence during an active production outage, capture a raw incident timeline, track response actions and system signals, and generate structured handoff notes to support later root cause analysis.

When should I start root cause analysis during an incident response?

Root cause analysis should only begin after you stabilize the incident, validate that mitigation actions reduced impact, and completely separate the repair process from your urgent outage mitigation efforts.

Can I use this for data-risk events and alert fires, or only full outages?

Yes, you can use this for data-risk events, alert fires, and production degradations, as it applies to any urgent operational failure requiring rapid stabilization, incident scoping, and structured stakeholder communication.

How do I scope the blast radius of a production degradation?

To scope the blast radius of a production degradation, confirm whether impact improved after mitigation, then systematically narrow the affected scope across users, regions, services, and workflows.