agency-sre-site-reliability-engineer

Implement SLO management, observability, and chaos engineering for production systems.

1|Updated May 5, 2026
One-click install
npx skills add https://github.com/bomberoxenviosdosruedas/01EnviosDosRueda --skill agency-sre-site-reliability-engineer-bomberoxenviosdosruedas
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: agency-sre-site-reliability-engineer
Source: https://github.com/bomberoxenviosdosruedas/01EnviosDosRueda/tree/main/.agents/workflows/agency-sre-site-reliability-engineer
Command: npx skills add https://github.com/bomberoxenviosdosruedas/01EnviosDosRueda --skill agency-sre-site-reliability-engineer-bomberoxenviosdosruedas

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires srelib, monitoring_tool, alerting_service, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides a site reliability engineer expert specializing in service reliability, observability, and chaos engineering for production systems at scale, helping to manage risks and ensure reliability through SLOs, error budgets, and automated tools.

Core Features & Use Cases

  • SLOs and Error Budget Management: Defines and monitors Service Level Objectives and error budgets to ensure reliability.
  • Observability: Implements logging, metrics, and tracing for rapid issue diagnosis and system health monitoring.
  • Toil Reduction: Automates routine tasks to allow engineers to focus on critical work.
  • Chaos Engineering: Conducts experiments to improve system resilience.
  • Capacity Planning: Optimizes resource allocation based on data-driven insights.
  • Use Case: An SRE could use this skill to define and enforce a new SLO for a critical service, set up alerts for performance anomalies, or create automated scripts to perform daily system checks.

Quick Start

Use the agency-sre-site-reliability-engineer skill to set up a new SLO for system availability with a 99.99% target.

Frequently Asked Questions about agency-sre-site-reliability-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define and monitor SLOs and error budgets for large-scale production systems?

You can define and monitor SLOs and error budgets for large-scale production systems by implementing Site Reliability Engineering principles, using integrated monitoring and alerting tools to track service reliability and automate routine checks.

How does chaos engineering improve system resilience in production environments?

Chaos engineering improves system resilience in production environments by conducting controlled experiments that identify weaknesses, allowing engineers to proactively address failures and optimize capacity planning based on data-driven insights.

What is the best way to set up observability for rapid issue diagnosis?

The best way to set up observability for rapid issue diagnosis is by implementing comprehensive logging, metrics, and tracing, which enables continuous system health monitoring and quick identification of performance anomalies.

Can I use this SRE approach with my existing monitoring_tool and alerting_service?

Yes, you can use this SRE approach with existing monitoring and alerting tools, as the methodology explicitly depends on these services to track SLO performance, manage error budgets, and trigger automated alerts for anomalies.

How do I automate toil reduction tasks for critical services?

You can automate toil reduction tasks for critical services by creating automated scripts to perform daily system checks, freeing engineers to focus on complex reliability work rather than routine operational duties.

Does SLO management work for services requiring 99.99% availability targets?

SLO management works for services requiring 99.99% availability targets by strictly defining and enforcing Service Level Objectives, tracking error budgets, and optimizing resource allocation to ensure high reliability at scale.