ai-sre-incident-response

Define AI incident classes, severity frameworks, and response playbooks for LLM outages.

46|4|Updated Jan 27, 2026
One-click install
npx skills add https://github.com/BagelHole/DevOps-Security-Agent-Skills --skill ai-sre-incident-response
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ai-sre-incident-response
Source: https://github.com/BagelHole/DevOps-Security-Agent-Skills/tree/main/devops/ai/ai-sre-incident-response
Command: npx skills add https://github.com/BagelHole/DevOps-Security-Agent-Skills --skill ai-sre-incident-response

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill establishes robust Site Reliability Engineering (SRE) practices tailored for AI systems, addressing unique incident classes like model outages, quality degradation, safety regressions, and cost overruns.

Core Features & Use Cases

  • AI Incident Classification: Defines and categorizes incidents specific to AI services (Availability, Quality, Safety, Cost).
  • Severity Framework: Provides a clear severity scale (SEV1-SEV3) for AI-related incidents.
  • Golden Signals: Identifies key metrics for monitoring AI service health.
  • Response Playbooks: Outlines specific actions for common AI incidents like model outages, quality regressions, and cost spikes.
  • Postmortem Requirements: Details essential elements for thorough incident reviews.

Quick Start

Apply SRE principles to AI systems by defining incident classes and response playbooks for AI outages and quality regressions.

Frequently Asked Questions about ai-sre-incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I establish incident response practices for LLM outages and quality degradation?

AI SRE incident response practices define incident classes, severity frameworks, and golden signals to address LLM outages and quality degradation. They provide specific response playbooks and postmortem requirements to mitigate availability, quality, safety, and cost incidents.

What are the key golden signals for monitoring AI service health?

Golden signals for monitoring AI service health identify key metrics across availability, quality, safety, and cost dimensions. These metrics help detect model outages, quality regressions, safety issues, and cost overruns specific to LLM services.

How do I classify AI incident severity levels for degraded model performance?

AI incident severity is classified using a clear SEV1-SEV3 severity framework tailored for AI services. This framework categorizes incidents across availability, quality, safety, and cost dimensions to prioritize response actions for degraded model performance.

Can I use SRE playbooks to mitigate runaway LLM API costs?

Yes, SRE response playbooks outline specific actions for cost incidents like runaway LLM API costs. They define mitigation strategies and monitoring metrics to detect and resolve cost overruns within AI services.

What should be included in an AI incident postmortem for safety regressions?

AI incident postmortems detail essential elements for thorough incident reviews following safety regressions. They require documenting specific mitigation strategies, monitoring metrics, and response actions taken during the safety incident.

When do I need a dedicated SRE framework for AI systems instead of traditional DevOps?

You need a dedicated AI SRE framework when addressing unique AI incident classes like model outages, safety regressions, and degraded quality that traditional DevOps practices do not cover. It establishes golden signals and playbooks specific to LLM services.