sre-engineer

Define SLIs/SLOs, manage incidents, and implement chaos engineering practices.

Updated Apr 25, 2026
One-click install
npx skills add https://github.com/Serg28/demosite --skill sre-engineer-serg28
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre-engineer
Source: https://github.com/Serg28/demosite/tree/main/.agents/skills/sre-engineer
Command: npx skills add https://github.com/Serg28/demosite --skill sre-engineer-serg28

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides comprehensive guidance for establishing and maintaining high reliability standards in production systems, ensuring minimal downtime and optimal performance.

Core Features & Use Cases

  • Define SLIs/SLOs: Help teams set measurable service level objectives aligned with user expectations.
  • Incident Management Procedures: Offer structured runbooks for detecting, responding to, and analyzing incidents efficiently.
  • Automation & Monitoring: Assist in developing scripts and configurations for proactive alerts, automated toil reduction, and resilience testing.
  • Use Case: An SRE team seeks to improve site reliability by implementing error budgets, chaos engineering experiments, and blameless postmortems across multiple services.

Quick Start

Describe your current system reliability goals, then review the provided scripts and documentation to start establishing SLIs, setting SLOs, and automating incident responses today.

Frequently Asked Questions about sre-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLIs and SLOs to improve service reliability?

To define SLIs and SLOs for service reliability, you establish measurable service level objectives aligned with user expectations. This Skill provides methodologies and scripts to set error budgets and track indicators across multiple services.

What is chaos engineering and how does it test infrastructure resilience?

Chaos engineering is the practice of running controlled experiments to test infrastructure resilience against failures. This Skill provides automation scripts to implement these experiments, ensuring your scalable systems maintain high reliability.

How do I create structured runbooks for incident management?

To create structured runbooks for incident management, you use procedures for detecting, responding to, and analyzing incidents efficiently. This Skill offers automation scripts to reduce toil and enable blameless postmortems.

Can I automate alerts and monitoring for production systems without external dependencies?

Yes, you can automate alerts and monitoring for production systems without external dependencies. This Skill provides standalone scripts and configurations for proactive alerting and automated toil reduction.

What is the best way to manage error budgets across multiple services?

The best way to manage error budgets across multiple services is to implement SLOs aligned with user expectations. This Skill provides the methodologies and automation scripts to track and enforce these budgets effectively.

When should I implement blameless postmortems for incident management?

You should implement blameless postmortems for incident management after resolving an incident to analyze what happened without pointing fingers. This Skill supports this practice to maintain high reliability standards and optimal performance.