Alerting & Incident Management

Implement alerting and incident response workflows with Prometheus, Alertmanager, and PagerDuty.

4|1|Updated Dec 17, 2025
One-click install
npx skills add https://github.com/lapc506/flutter-agentic-boilerplate --skill alerting-incident-management-lapc506
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: Alerting & Incident Management
Source: https://github.com/lapc506/flutter-agentic-boilerplate/tree/main/skills/system-reliability-engineering/alerting-incident-management
Command: npx skills add https://github.com/lapc506/flutter-agentic-boilerplate --skill alerting-incident-management-lapc506

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill addresses the critical need for maintaining reliable services by establishing effective alerting mechanisms and streamlined incident response processes.

Core Features & Use Cases

  • Effective Alerting: Design and implement alerts that accurately reflect service health and potential issues.
  • Incident Management: Define clear workflows for on-call rotations, incident response, and resolution.
  • Runbooks: Create actionable guides for diagnosing and resolving common incidents.
  • Use Case: When a critical service experiences high latency or becomes unavailable, this Skill ensures the right on-call engineer is notified immediately via PagerDuty, provided with a detailed runbook to diagnose the issue, and guided through the resolution process.

Quick Start

Configure alerting with PagerDuty and create runbooks for critical services.

Frequently Asked Questions about Alerting & Incident Management

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I set up Prometheus alerting rules for high service latency?

To set up Prometheus alerting for high service latency, you design alerts that accurately reflect service health and potential issues. This approach ensures proactive issue detection to maintain high availability and meet service level agreements.

What is the best way to structure on-call rotations for incident response?

The best way to structure on-call rotations involves defining clear workflows for incident response and resolution. This guarantees the right on-call engineer is notified immediately via PagerDuty when a critical service experiences issues.

How do I create actionable runbooks for diagnosing production incidents?

You create actionable runbooks by developing detailed guides for diagnosing and resolving common incidents. These runbooks guide on-call engineers through the resolution process during critical service disruptions.

Can I integrate Alertmanager with PagerDuty for immediate engineer notifications?

Yes, you can integrate Alertmanager with PagerDuty to ensure the right on-call engineer is notified immediately. This integration facilitates rapid, organized incident response to maintain reliable services.

Why do I need comprehensive incident management for maintaining service reliability?

You need comprehensive incident management to establish effective alerting mechanisms and streamlined incident response processes. It addresses the critical need for maintaining reliable services and meeting service level agreements.