sre-engineer

Create structured SRE playbooks for Kubernetes microservices with SLOs, alerts, and runbooks.

9|1|Updated Mar 21, 2026
One-click install
npx skills add https://github.com/e-t-y-b/etyb-skills --skill sre-engineer-e-t-y-b
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre-engineer
Source: https://github.com/e-t-y-b/etyb-skills/tree/main/skills/sre-engineer
Command: npx skills add https://github.com/e-t-y-b/etyb-skills --skill sre-engineer-e-t-y-b

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill provides structured SRE guidance to design, implement, and operate reliable software systems across monitoring, incident response, capacity planning, logging, tracing, and chaos engineering.

Core Features & Use Cases

  • Observability foundations: define SLOs/SLIs, burn-rate alerts, dashboards, and correlation of metrics/logs/traces.
  • Incident lifecycle: on-call design, runbooks, postmortems, escalation, and war-room playbooks.
  • Resilience & capacity planning: chaos engineering, auto-scaling patterns, FinOps, and capacity modeling for scalable services.

Quick Start

Describe your reliability challenge and I will output a tailored SRE playbook and references.

Frequently Asked Questions about sre-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLOs and SLI burn-rate alerts for a Kubernetes microservices architecture?

To define SLOs and burn-rate alerts for Kubernetes microservices, this Skill generates a structured SRE playbook that establishes service level indicators, configures error budget tracking, and sets up multi-window burn-rate alerting rules to maintain production reliability targets.

What is the best way to structure incident response runbooks and postmortems for multi-team ownership?

The best way to structure incident response runbooks and postmortems for multi-team ownership is to use a unified SRE playbook that defines on-call rotations, escalation paths, war-room procedures, and blameless postmortem templates to coordinate resolution across distributed teams.

How does chaos engineering fit into capacity planning and auto-scaling patterns for reliable services?

Chaos engineering integrates into capacity planning by proactively validating auto-scaling patterns and resilience mechanisms under failure conditions, allowing teams to model capacity accurately and ensure scalable services meet production reliability targets before incidents occur.

Can I use this SRE playbook to implement FinOps practices and observability foundations together?

Yes, you can use this SRE playbook to implement FinOps practices alongside observability foundations, as it unifies capacity modeling and cost optimization with metrics, logging, and tracing correlation to provide comprehensive reliability and financial operations guidance.

Do I need a specific monitoring stack to apply these SLO definitions and incident management playbooks?

No specific monitoring stack is required to apply these SLO definitions and incident management playbooks, as the Skill provides vendor-agnostic structured guidance for designing reliability frameworks that can be implemented with various observability and alerting tools.

Why does my burn-rate alerting strategy fail to prevent SLO violations during peak traffic spikes?

Your burn-rate alerting strategy may fail during peak traffic spikes if it lacks proper multi-window configurations or accurate capacity modeling, which this Skill addresses by structuring alert thresholds, auto-scaling patterns, and chaos testing to validate resilience under load.