incident-response

Automate incident detection, tracking, escalation, resolution, and post-mortem analysis via Datadog API.

Updated Jan 14, 2022
One-click install
npx skills add https://github.com/alexmarucci/dotfiles --skill incident-response-alexmarucci
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: incident-response
Source: https://github.com/alexmarucci/dotfiles/tree/main/claude/skills/incident-response
Command: npx skills add https://github.com/alexmarucci/dotfiles --skill incident-response-alexmarucci

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires pup, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides a comprehensive solution for incident response, simplifying on-call management, incident tracking, and case management for service reliability.

Core Features & Use Cases

  • On-Call Management: Handle on-call schedules, escalations, and pagers for effective alerting and management.
  • Incident Management: Create, track, and resolve incidents, with full lifecycle management.
  • Case Management: Create, update, and manage cases for incidents, with full history tracking and project management.
  • Use Case: When an incident occurs, the Skill can automate the creation of an incident record, assign the case to an incident commander, escalate alerts, track the status of the incident, and finally close the case post-resolution.

Quick Start

To activate incident response, type "start incident response" in your context.

Frequently Asked Questions about incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate incident tracking and on-call scheduling for service reliability?

Incident tracking and on-call scheduling are automated by creating incident records, assigning incident commanders, escalating alerts, and resolving cases with full lifecycle history tracking. This workflow supports production-grade service reliability management.

What is the best way to handle incident escalation and post-mortem analysis?

Incident escalation and post-mortem analysis are handled through a comprehensive workflow that automates detection, tracks status, manages escalations via pagers, and closes cases post-resolution while retaining full case history for review.

Does incident response management require Datadog API access to function?

Yes, Datadog API access is required to enable the incident response workflow. This integration supports automated detection, tracking, and escalation of incidents within your service reliability environment.

Can I manage full case histories and incident lifecycles for production-grade systems?

You can manage full incident lifecycles and case histories for production-grade systems. The workflow creates, updates, and resolves cases while tracking complete incident records from initial alert through post-resolution closure.

How do I start an incident response workflow when an alert triggers?

To start an incident response workflow, type "start incident response" in your context. This activates the automated creation of incident records, assigns an incident commander, and initiates alert escalation and tracking.