incident-response

Identify production incidents and plan rollback with severity and blast-radius assessment.

16|Updated Apr 30, 2026
One-click install
npx skills add https://github.com/JCE-Joshhh77/JCE-Opencode-Tools --skill incident-response-jce-joshhh77
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-response
Source: https://github.com/JCE-Joshhh77/JCE-Opencode-Tools/tree/main/config/skills/incident-response
Command: npx skills add https://github.com/JCE-Joshhh77/JCE-Opencode-Tools --skill incident-response-jce-joshhh77

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Production incidents are time-sensitive and complex; this skill provides a structured, repeatable approach to triage, mitigate, and document outages to restore services quickly and minimize blast radius.

Core Features & Use Cases

  • Structured triage workflow with quick confirmation and scope steps
  • Severity classification, blast-radius assessment, and rollback planning
  • Rollback safety checks and post-mortem templates for blameless investigations
  • Guidance for communication during incidents and post-stabilization verification
  • Integration with existing incident workflows and runbooks

Quick Start

Describe the incident and activate the triage workflow to stabilize the system, communicate status, and capture a blameless post-mortem.

Frequently Asked Questions about incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I triage a production outage to assess severity and blast radius?

Triage a production outage by identifying the incident scope, classifying severity, and evaluating the blast radius to enable rapid stabilization and targeted rollback planning. This structured approach minimizes user impact and restores services safely.

What is a blameless post-mortem and when do I need one after an incident?

A blameless post-mortem is a structured root-cause analysis template used after stabilizing an incident. You need one following any outage, deploy failure, or user-impact event to capture repeatable, safe mitigations and drive future prevention.

How do I plan a safe rollback after a deploy failure?

Plan a safe rollback after a deploy failure by applying rollback safety checks and scoped triage steps to verify the deployment state. This ensures mitigations are repeatable and prevents further blast radius expansion during restoration.

Can I integrate incident response workflows with my existing runbooks?

You can integrate incident response workflows with existing runbooks through specified integration points. The skill guides communication during incidents and connects with your established workflow to ensure safe, repeatable mitigations and post-stabilization verification.

What is the best way to stabilize outages fast and minimize user impact?

The best way to stabilize outages fast is applying a structured triage workflow that assesses blast radius and plans rollbacks. Quick confirmation, scope identification, and severity classification enable rapid stabilization and minimize the impact of deploy failures.