sre-engineer

Define SLOs, error budgets, monitoring, and automation scripts for production systems.

1|Updated Jan 19, 2026
One-click install
npx skills add https://github.com/camelranchentertainment/Booking-Platform --skill sre-engineer-camelranchentertainment
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre-engineer
Source: https://github.com/camelranchentertainment/Booking-Platform/tree/main/.claude/skills/sre-engineer
Command: npx skills add https://github.com/camelranchentertainment/Booking-Platform --skill sre-engineer-camelranchentertainment

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Defines how to establish reliability for production systems by setting SLOs/SLIs, creating error budgets, designing incident response, modeling capacity, and delivering monitoring configurations and automation scripts.

Core Features & Use Cases

  • Define quantitative SLOs/SLIs and error budgets with burn-rate monitoring.
  • Create runbooks, blameless postmortems, and incident response procedures; establish golden signals monitoring.
  • Produce automation scripts and capacity planning models for scalable reliability.

Quick Start

Configure SLOs, error budgets, incident response procedures, and monitoring setups for a production service.

Frequently Asked Questions about sre-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLOs and error budgets for production services?

Define SLOs and error budgets by establishing quantitative SLIs and configuring burn-rate monitoring to track reliability thresholds for production services.

What is the best way to set up incident management and blameless postmortems?

Set up incident management by creating runbooks and incident response procedures, then conduct blameless postmortems to analyze failures without assigning individual fault.

How do I configure golden signals monitoring for scalable services?

Configure golden signals monitoring by tracking latency, traffic, errors, and saturation metrics to establish measurable SLIs across scalable service infrastructure.

Can I automate runbooks and capacity planning for site reliability engineering?

Yes, you can automate runbooks and generate capacity planning models to reduce operational toil and deliver quantitative reliability automation scripts for scalable services.

When do I need chaos engineering for site reliability engineering?

Apply chaos engineering when you need to proactively test production system resilience by running controlled experiments that validate incident response procedures and reliability thresholds.