agency-sre-site-reliability-engineer

Define SLOs and automate toil for production reliability.

Updated Mar 22, 2026
One-click install
npx skills add https://github.com/jay6697117/agency-agents-antigravity --skill agency-sre-site-reliability-engineer
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: agency-sre-site-reliability-engineer
Source: https://github.com/jay6697117/agency-agents-antigravity/tree/main/.agents/skills/agency-sre-site-reliability-engineer
Command: npx skills add https://github.com/jay6697117/agency-agents-antigravity --skill agency-sre-site-reliability-engineer

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Production systems often suffer from unclear reliability targets, manual toil, and slow incident response. This Skill codifies reliability as a feature by defining SLOs, improving observability, and automating repetitive operational work.

Core Features & Use Cases

  • Define SLOs and error budgets aligned with user experience
  • Build metrics, logs, and traces to answer "why is this failing?" quickly
  • Automate toil reduction with repeatable recovery and incident playbooks
  • Enable chaos engineering experiments to proactively find weaknesses
  • Capacity planning based on data-driven usage and load

Quick Start

Define SLOs for your critical services and implement automated toil reduction workflows to improve reliability across your production stack.

Frequently Asked Questions about agency-sre-site-reliability-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLOs and error budgets aligned with user experience?

To define SLOs, you establish measurable reliability targets tied directly to user experience. This Skill configures an SLO framework that sets clear service level objectives and tracks error budgets, transforming reliability into a measurable budget for your production services.

What's the best way to automate toil reduction for incident response?

Automating toil reduction involves building repeatable recovery workflows and automated incident playbooks. This Skill enables you to automate repetitive operational work, significantly speeding up incident response times across multiple services.

How does observability guidance help answer why production systems are failing?

Observability guidance helps answer why systems are failing by integrating metrics, logs, and traces. This Skill builds comprehensive observability into your production stack so you can quickly diagnose failures and understand system behavior during incidents.

Can I use chaos engineering experiments for capacity planning in large-scale production systems?

Yes, you can use chaos engineering experiments alongside capacity planning for large-scale production systems. This Skill enables proactive chaos engineering to find weaknesses and applies data-driven usage and load analysis for effective capacity planning across services.

Do I need progressive rollout governance to improve production reliability?

You need progressive rollout governance to safely improve production reliability and prevent widespread incidents. This Skill provides governance for progressive rollouts, ensuring that new releases do not exhaust your error budgets or compromise service stability.