operations-incident

Triage and mitigate production incidents using a structured playbook.

Updated Jan 5, 2025
One-click install
npx skills add https://github.com/pkuppens/pkuppens --skill operations-incident
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: operations-incident
Source: https://github.com/pkuppens/pkuppens/tree/main/skills/operations/operations-incident
Command: npx skills add https://github.com/pkuppens/pkuppens --skill operations-incident

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Incidents in production can cause downtime and confusion; this Skill provides a repeatable, fast-response workflow to triage, mitigate, and document incidents, reducing response time and improving incident reports.

Core Features & Use Cases

  • Triage steps, mitigation actions, status communications, and post-incident reporting.
  • Use cases include degraded services, alert fires, and root-cause investigations with structured timelines.

Quick Start

Follow the incident playbook to triage, mitigate, and compose a post-incident timeline.

Frequently Asked Questions about operations-incident

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I triage a production incident when an alert fires?

To triage a production incident when an alert fires, follow a structured playbook that identifies the issue across services and deployment environments, then applies mitigation actions to reduce downtime.

What is the best way to document a post-incident timeline for root-cause analysis?

The best way to document a post-incident timeline for root-cause analysis is to enforce a structured post-incident reporting workflow that captures verification steps and mitigation actions.

Can I use this incident response workflow across different deployment environments and data stores?

Yes, you can use this incident response workflow across different deployment environments and data stores because the playbook applies to outages and alert fires spanning multiple services.

What steps should I follow to mitigate degraded services during an outage?

To mitigate degraded services during an outage, follow the incident playbook to identify the root-cause, apply mitigation actions, execute verification steps, and compose a post-incident timeline.

Why do I need a structured playbook for on-call incident response?

You need a structured playbook for on-call incident response because production incidents cause downtime and confusion, and a repeatable fast-response workflow reduces response time and improves reports.

Does this incident triage process include status communications and post-incident reporting?

Yes, this incident triage process includes status communications and post-incident reporting as core features to manage degraded services, alert fires, and root-cause investigations.