ops-incident-response

Coordinate production incident phases from detection to post-incident review.

Updated Mar 21, 2026
One-click install
npx skills add https://github.com/LucasMalessa/TheRing --skill ops-incident-response-lucasmalessa
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ops-incident-response
Source: https://github.com/LucasMalessa/TheRing/tree/main/.archive/ops-team/skills/ops-incident-response
Command: npx skills add https://github.com/LucasMalessa/TheRing --skill ops-incident-response-lucasmalessa

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

The Incident Response Workflow provides a structured, repeatable process to detect, declare, triage, mitigate, and post-mortem production incidents, reducing mean time to recovery and improving reliability.

Core Features & Use Cases

  • Phase-driven playbook covering Detection, Declaration, Triage, Mitigation, Resolution, and Post-Incident reviews.
  • Templates and owner assignments for incident channels, RCA, and retrospective actions to ensure blameless learning.
  • Use Case: When a Sev1 outage occurs, the team follows the workflow to coordinate responders, document impact, coordinate remediation, and schedule RCA.

Quick Start

Trigger the incident response workflow for a live incident and follow the six phases to manage the event from detection to post-mortem.

Frequently Asked Questions about ops-incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I coordinate incident response for a production outage from detection to post-mortem?

Incident response coordination follows a six-phase workflow covering detection, declaration, triage, mitigation, resolution, and post-incident review. Each phase provides structured steps, role assignments, and runbook templates to standardize the process and reduce mean time to recovery.

What is a structured incident response workflow and when do I need one?

A structured incident response workflow standardizes how teams detect, triage, mitigate, and review production incidents. You need one for Sev1, Sev2, or Sev3 outages, degradations, and security events to ensure repeatable coordination, blameless post-mortems, and root cause analysis across services and teams.

How do I run a blameless post-mortem after resolving a Sev1 incident?

Post-mortem reviews use provided templates and owner assignments to document impact, coordinate remediation, and schedule root cause analysis. The workflow ensures blameless learning by standardizing retrospective actions and assigning clear ownership for follow-up items after the incident is resolved.

Can I use this incident management workflow for both service degradations and security events?

Yes, the incident management workflow applies to Sev1, Sev2, and Sev3 incidents including outages, service degradations, and security events. It standardizes response across multiple services and teams using phase-driven playbooks, role assignments, and runbook templates regardless of incident type.

What is the best way to assign roles and manage an incident channel during a live outage?

Role assignments and incident channel templates are provided during the declaration and triage phases of the workflow. This ensures responders are coordinated, impact is documented, and remediation efforts are structured across teams throughout the live incident.