sre-engineering

Coordinate production incident response and post-mortem analysis.

2|Updated Mar 16, 2026
One-click install
npx skills add https://github.com/chicongst/agent-skills-installer --skill sre-engineering
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre-engineering
Source: https://github.com/chicongst/agent-skills-installer/tree/main/skills/sre-engineering
Command: npx skills add https://github.com/chicongst/agent-skills-installer --skill sre-engineering

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Use when handling production incidents, reliability concerns, on-call triage, infrastructure issues, monitoring, alerting, and post-mortem analysis.

Core Features & Use Cases

  • Incident triage and coordination across on-call teams to rapidly assess impact and blast radius.
  • Structured response framework with parallel investigation tracks, time-boxed actions, and regular stakeholder communications.
  • Post-incident analysis, reporting, and continuous improvement through documented RCA and preventive measures.

Quick Start

Describe and execute an incident response plan for a production outage, including roles, timelines, and mitigation steps.

Frequently Asked Questions about sre-engineering

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I coordinate a production incident response across multiple on-call teams?

Production incident response coordination requires a structured framework defining severity levels, parallel investigation tracks, time-boxed actions, and stakeholder communication templates to assess blast radius and mitigate impact across on-call teams.

What is the best way to structure a post-incident review and RCA?

A post-incident review and RCA should document structured timelines, escalation paths, decision governance, and preventive measures to drive continuous improvement and reliability across affected services and systems.

How does severity level classification work during an infrastructure outage?

Severity level classification during an infrastructure outage applies an incident response framework defining escalation paths and decision governance to triage impact, coordinate on-call responses, and structure stakeholder communications.

Can I use this incident response framework for monitoring alerts and reliability concerns, not just full outages?

Yes, this incident response framework applies to monitoring alerts, reliability concerns, on-call triage, and infrastructure issues across services, providing structured timelines and communication templates beyond full production outages.

What do I need to set up an incident response plan for a production outage?

Setting up an incident response plan for a production outage requires defining incident roles, structured timelines, mitigation steps, escalation paths, and communication templates to execute coordinated response and post-mortem analysis.

When should I not use a structured incident response playbook?

A structured incident response playbook should not be used for routine operational tasks lacking production impact, as its severity levels, escalation paths, and governance are designed for outages, infrastructure incidents, and reliability concerns.