site-reliability-engineer

Define SLI/SLO targets and burn-rate policies for production reliability.

7|1|Updated May 19, 2026
One-click install
npx skills add https://github.com/daemon-blockint-tech/Agentic-Enteprises-Skill --skill site-reliability-engineer-daemon-blockint-tech
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: site-reliability-engineer
Source: https://github.com/daemon-blockint-tech/Agentic-Enteprises-Skill/tree/main/site-reliability-engineer
Command: npx skills add https://github.com/daemon-blockint-tech/Agentic-Enteprises-Skill --skill site-reliability-engineer-daemon-blockint-tech

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Site reliability engineering guidance to define SLOs, manage error budgets, and improve production readiness across services, reducing downtime and toil.

Core Features & Use Cases

  • Define SLIs/SLOs and error budgets; configure burn-rate alerts and dashboards.
  • Conduct production readiness reviews and capacity planning; map dependencies and failure modes.
  • Lead incident mitigation and release reliability activities (PRR, canaries, chaos testing) to protect customer impact.

Quick Start

Define an SLI/SLO for a service and draft a burn-rate policy in preparation for a canary release.

Frequently Asked Questions about site-reliability-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLI and SLO targets for a multi-service system?

To define SLI and SLO targets for a multi-service system, establish service level indicators and reliability objectives, then generate SLO documents and burn-rate policies to govern production reliability and track error budgets.

What is an error budget burn-rate policy and when do I need it?

An error budget burn-rate policy tracks how quickly your service consumes its allowed failure threshold against SLOs. You need it to configure alerts and dashboards that trigger incident response before customer impact occurs.

How do I conduct a production readiness review for a canary release?

To conduct a production readiness review for a canary release, apply SLO governance rules across capacity planning and dependency mapping, then generate PRR checklists to validate release reliability and failure mode mitigation.

Can I use SRE practices for incident response and on-call mitigation?

Yes, you can use SRE practices for incident response by creating dependency maps and runbooks tailored for on-call mitigation. This protects customer impact by applying established SLO rules during active incident mitigation.

What is the best way to plan capacity and map failure modes across services?

The best way to plan capacity and map failure modes across services is to apply defined SLO targets to dependency maps, ensuring your multi-service systems maintain production readiness and reduce ongoing operational toil.