sre-engineer

Define SLOs, error budgets, and incident response procedures for production systems.

Updated Jan 9, 2026
One-click install
npx skills add https://github.com/dieu-donnee/luxtrax --skill sre-engineer-dieu-donnee
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre-engineer
Source: https://github.com/dieu-donnee/luxtrax/tree/main/.agent/skills/sre-engineer
Command: npx skills add https://github.com/dieu-donnee/luxtrax --skill sre-engineer-dieu-donnee

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

The SRE Engineer skill guides teams in designing and operationalizing reliability practices, including SLO/SLI definitions, error budget policies, incident response playbooks, capacity planning, and monitoring configurations.

Core Features & Use Cases

  • Define and align SLOs/SLIs to user impact across services.
  • Create and manage error budget policies, burn rate alerts, and runbooks.
  • Design incident response procedures, postmortems, capacity planning models, and monitoring setups.

Quick Start

Provide a concise SLO plan and initial monitoring outline for your production service.

Frequently Asked Questions about sre-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define and align SLOs and SLIs to measure user impact across services?

To define SLOs and SLIs, you establish quantifiable reliability targets mapped directly to user impact across distributed services. This Skill generates SLO plans and initial monitoring outlines, translating service-level indicators into actionable error budget policies.

What is an error budget policy and how does it guide release decisions?

An error budget policy dictates the permissible failure rate within an SLO framework. This Skill designs burn-rate alerts and error budget management rules to balance feature velocity against reliability requirements for production systems.

How do I set up incident response procedures and blameless postmortems for distributed services?

Setting up incident response procedures involves creating runbooks and postmortem templates for production incidents. This Skill designs blameless postmortem workflows and incident management playbooks tailored for enterprise-grade software teams.

Can I use this for capacity planning and monitoring configurations in enterprise environments?

Yes, this Skill targets enterprise-grade software teams by generating capacity planning models and monitoring configurations. It enforces end-to-end reliability requirements including automation templates and burn-rate alerts for distributed services.

What is the best way to start operationalizing reliability practices for a new production service?

The best way to start operationalizing reliability practices is to provide a concise SLO plan and initial monitoring outline. This Skill guides teams in building incident response playbooks and aligning SLI measurements to user impact from day one.

SRE Engineer: does it support generating runbooks and automation templates for incident management?

Yes, the SRE Engineer Skill supports generating runbooks and automation templates for incident management. It enforces end-to-end reliability requirements by designing incident response procedures and blameless postmortems for distributed services.