sre

Define SLO targets, burn-rate alerts, and incident runbooks for production systems.

Updated Apr 10, 2026
One-click install
npx skills add https://github.com/Exia-thd/Digital-Nervous --skill sre-exia-thd
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre
Source: https://github.com/Exia-thd/Digital-Nervous/tree/main/skills/sre
Command: npx skills add https://github.com/Exia-thd/Digital-Nervous --skill sre-exia-thd

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Site reliability engineering provides a framework to make production systems reliable through SLOs, monitoring, incident runbooks, chaos testing, and capacity planning.

Core Features & Use Cases

  • Define and enforce service-level objectives (SLOs) and error budgets across teams.
  • Build incident response procedures, runbooks, and disaster-recovery artifacts to shorten outages.
  • Validate resilience with chaos experiments and capacity planning to guide scaling decisions.

Quick Start

Define SLOs for your services and implement foundational runbooks to start improving reliability.

Frequently Asked Questions about sre

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLOs and error budgets for my production services?

To define SLOs, you establish service-level objective targets using existing infrastructure and monitoring data as inputs. This generates concrete SLO definitions, burn-rate alerts, and dashboards to track error budgets across your services.

What is the best way to set up incident management runbooks and disaster recovery docs?

Setting up incident management runbooks involves defining response procedures and readiness checks as inputs. This generates actionable runbooks and disaster recovery docs to systematically shorten outages and guide recovery operations.

How does chaos engineering validate resilience and guide capacity planning?

Chaos engineering validates resilience by executing chaos experiments against your architecture. Combined with capacity planning, this process tests failure scenarios and outputs actionable scaling decisions to guide infrastructure growth.

Can I use this SRE framework with my existing monitoring and infrastructure setup?

Yes, you can use this SRE framework with existing monitoring. It takes your current infrastructure data and architecture docs as inputs to generate SLO definitions, burn-rate alerts, and dashboards without requiring external dependencies.

What do I need to start implementing site reliability engineering for my systems?

To start implementing site reliability engineering, you provide infrastructure monitoring data, architecture docs, and readiness checks. These inputs enable you to quickly define SLOs and build foundational runbooks for your production systems.