sre-engineer

Define SLI/SLO targets and error budgets for service reliability.

Updated Feb 11, 2026
One-click install
npx skills add https://github.com/lamb92009/claude-skills --skill sre-engineer-lamb92009
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre-engineer
Source: https://github.com/lamb92009/claude-skills/tree/main/sre-engineer
Command: npx skills add https://github.com/lamb92009/claude-skills --skill sre-engineer-lamb92009

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This skill helps define meaningful SLI/SLO targets, establish robust error budgets, and orchestrate reliable system practices across incidents, chaos experiments, toil reduction, capacity planning, and on-call workflows.

Core Features & Use Cases

  • Define and govern SLI/SLO targets with error budgets to quantify reliability for services.
  • Implement monitoring, postmortems, and runbooks; automate toil reduction and resilience workflows.
  • Apply to incident management, chaos engineering, capacity planning, and on-call reliability improvements.

Quick Start

Describe your service reliability goals and run through an SRE workflow to define SLIs/SLOs, error budgets, and incident response runbooks.

Frequently Asked Questions about sre-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLI and SLO targets with error budgets for my services?

To define SLI and SLO targets, you quantify service reliability by establishing error budgets that dictate acceptable failure rates. This approach governs incident management and monitoring by providing measurable thresholds for system availability.

What is the best way to set up incident management and postmortems using SRE practices?

Incident management and postmortems are set up by applying SRE workflows that document runbooks and golden signals. This process orchestrates reliable system practices by standardizing incident response and analyzing failures to reduce future outages.

Can I use chaos engineering and capacity planning to improve service reliability?

Chaos engineering and capacity planning are applied to test resilience and forecast system load. By integrating these practices with your defined SLOs, you validate system behavior under stress and ensure scalable reliability improvements.

How do I reduce toil and automate on-call workflows for scalable systems?

Toil reduction and on-call workflow automation are achieved by creating automation scripts that handle repetitive operational tasks. This minimizes manual intervention, allowing on-call practices to focus on scalable reliability workflows and incident response.

Does this SRE approach support creating monitoring runbooks and SLO governance?

This SRE approach supports creating SLO governance, monitoring runbooks, and postmortems. It helps establish robust error budgets and orchestrates reliable system practices across incidents, chaos experiments, and on-call workflows.