tech-devops-incident-responder

Generate structured runbooks and post-mortem templates for live incident triage.

Updated Jan 28, 2026
One-click install
npx skills add https://github.com/scanady/nexus-agents --skill tech-devops-incident-responder
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: tech-devops-incident-responder
Source: https://github.com/scanady/nexus-agents/tree/main/skills/tech-devops-incident-responder
Command: npx skills add https://github.com/scanady/nexus-agents --skill tech-devops-incident-responder

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill helps SRE teams triage, coordinate, and perform post-mortems quickly during live incidents, improving MTTR and reliability.

Core Features & Use Cases

  • Incident triage and severity classification to establish priorities
  • Runbook generation and on-call coordination during outages
  • Blameless post-mortem facilitation and root-cause analysis
  • SLO/error budget analysis guidance and reliability pattern suggestions
  • On-call workflow design and communication templates

Quick Start

Describe the incident details you are facing and request a ready-to-use runbook and post-mortem template.

Frequently Asked Questions about tech-devops-incident-responder

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I conduct a blameless post-mortem after a live incident?

A blameless post-mortem focuses on systemic root-cause analysis rather than individual fault. You facilitate it using structured templates to document timeline, contributing factors, and preventive actions, ensuring distributed systems reliability improves continuously.

How do I classify incident severity during an active outage?

Incident severity classification establishes priorities during an outage by evaluating impact on SLOs, error budgets, and user experience. Structured triage guidance helps categorize severity levels to coordinate on-call response and allocate resources effectively.

What is the best way to generate an on-call runbook for distributed systems?

The best way to generate an on-call runbook is to provide observability context and best-practice templates to produce actionable, step-by-step resolution guidance. This ensures on-call rotations have structured workflows during live outage events.

Can I use this for SLO and error budget analysis during incident response?

Yes, you can use it for SLO and error budget analysis during incident response. It provides guidance on evaluating reliability patterns and error budget consumption to inform triage decisions and prioritize remediation across distributed systems.

What observability data do I need for effective incident triage?

Effective incident triage requires observability context and incident data to generate actionable guidance. Supplying detailed outage metrics, logs, and best-practice runbook templates enables accurate severity classification and structured on-call coordination.