Alerting & Incident Management

Design alerting and incident management workflows for production services.

1|Updated Apr 28, 2024
One-click install
npx skills add https://github.com/HabitaNexus/monorepo --skill alerting-incident-management-habitanexus
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: Alerting & Incident Management
Source: https://github.com/HabitaNexus/monorepo/tree/main/skills/system-reliability-engineering/alerting-incident-management
Command: npx skills add https://github.com/HabitaNexus/monorepo --skill alerting-incident-management-habitanexus

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Alerting and incident management are essential to keep production services reliable. This skill helps teams design effective alerting, on-call rotations, and runbooks to reduce MTTR and improve incident handling.

Core Features & Use Cases

  • Design and implement multi-channel alerting and on-call routing
  • Create runbooks and incident response workflows for consistent actions
  • Integrate with PagerDuty, Opsgenie, Slack, and Email for timely notifications
  • Use post-incident reviews to close feedback loops and prevent recurrence

Quick Start

Configure an on-call rotation and runbook for a critical service and connect PagerDuty for incident notifications.

Frequently Asked Questions about Alerting & Incident Management

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I set up on-call rotations and alerting for production services?

Set up on-call rotations and alerting by configuring multi-channel alert rules, routing policies, and runbooks for production services. This orchestrates consistent incident notifications and reduces mean time to resolve across distributed systems.

What is incident management and when do I need runbooks for incident response?

Incident management is the orchestration of alert routing, on-call schedules, and structured response workflows. You need runbooks when managing distributed systems to ensure consistent actions during production incidents and reduce resolution time.

Does this incident response workflow integrate with PagerDuty and Opsgenie?

Yes, the incident response workflow integrates with PagerDuty, Opsgenie, Slack, and Email. This multi-channel notification setup ensures timely alert delivery to on-call engineers across your existing communication platforms.

What's the best way to configure alert routing and on-call policies for distributed systems?

The best way to configure alert routing and on-call policies is through end-to-end setup of alert rules, routing configurations, and on-call rotations. This orchestrates reliable incident response workflows tailored for distributed production systems.

How do I create post-incident reviews to close feedback loops after an on-call alert?

Create post-incident reviews by executing structured incident response workflows after on-call alerts. This closes feedback loops by documenting the incident handling process to prevent recurrence and improve future alert configurations.