sre-engineer

Define SLOs/SLIs, error budgets, and incident response automation for reliability.

16|Updated Apr 19, 2026
One-click install
npx skills add https://github.com/Marwan78888/Neuron-Cli --skill sre-engineer-marwan78888
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre-engineer
Source: https://github.com/Marwan78888/Neuron-Cli/tree/main/scratch/claude-skills-main/skills/sre-engineer
Command: npx skills add https://github.com/Marwan78888/Neuron-Cli --skill sre-engineer-marwan78888

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This skill provides a structured approach to defining SLOs/SLIs, budgeting error budgets, and automating incident response and monitoring to improve reliability across production systems.

Core Features & Use Cases

  • Define SLOs/SLIs with measurable targets and budget policies to quantify reliability
  • Design incident response playbooks, chaos engineering plans, and toil-reduction automation
  • Build capacity models and monitoring configurations to sustain reliability at scale

Quick Start

Define your first SLO and set up a basic monitoring dashboard to start measuring reliability.

Frequently Asked Questions about sre-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLOs and SLIs to measure system reliability?

To define SLOs and SLIs, establish measurable targets for service availability and performance, then apply error budget policies to quantify reliability across production systems. This skill provides a structured approach to defining these indicators and setting governance to enforce them.

What's the best way to automate incident response and reduce toil?

Automating incident response and reducing toil requires designing incident response playbooks and toil-reduction automation scripts. This framework builds runbooks and configurations to automate monitoring and sustain reliability at scale without manual intervention.

How does error budgeting work for managing service availability?

Error budgeting works by quantifying the allowable downtime or error rate against your defined SLOs. This skill structures error budget policies to measure and improve system availability, ensuring governance enforces reliability targets across your production environment.

Can I use this to build capacity models and monitoring configurations for scalable systems?

Yes, you can build capacity models and monitoring configurations for scalable systems. This skill designs reliability engineering practices that establish monitoring configurations and capacity planning to sustain system availability at scale.

Do I need existing runbooks and automation scripts to implement chaos engineering?

You do not need existing scripts to start, as this skill helps design chaos engineering plans and runbooks from scratch. It applies reliability engineering practices to establish the necessary automation scripts and governance to enforce reliability during chaos testing.

When should I establish measurable targets for production system reliability?

You should establish measurable targets for production system reliability before scaling operations. Defining SLOs and setting up basic monitoring dashboards early provides the baseline needed to quantify reliability and apply error budget policies effectively.