sre

Generate SRE artifacts including SLOs, runbooks, and incident processes for services.

47|11|Updated Mar 6, 2026
One-click install
npx skills add https://github.com/buiphucminhtam/forgewright --skill sre-buiphucminhtam
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre
Source: https://github.com/buiphucminhtam/forgewright/tree/main/skills/sre
Command: npx skills add https://github.com/buiphucminhtam/forgewright --skill sre-buiphucminhtam

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Production-grade reliability engineering for services, including SLOs, monitoring, incident runbooks, chaos engineering, and capacity planning to keep systems resilient.

Core Features & Use Cases

  • SLO definition and governance: Establish measurable reliability targets and ensure alignment with business expectations.
  • Phase-driven readiness: Guide through readiness reviews, monitoring, chaos experiments, and incident management across the lifecycle.
  • Runbooks and playbooks: Create incident response plans, escalation policies, and drill procedures tailored to each service.
  • Chaos engineering and resilience testing: Design and run controlled experiments to validate steady-state behavior and failure handling.
  • Capacity planning and scaling: Model load, plan capacity, and validate auto-scaling configurations.
  • Documentation and handoffs: Produce artifacts for platform, developers, and leadership.

Quick Start

Provide your service context and I will generate a complete SRE setup with readiness checks, SLOs, and runbooks.

Frequently Asked Questions about sre

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLOs and set up burn-rate alerts for my service?

To define SLOs and burn-rate alerts, this Skill generates measurable reliability targets aligned with business expectations and configures burn-rate alerts to notify you before exhausting your error budget. It provides concrete SLO definitions and governance templates tailored to your service context.

What is the best way to create incident runbooks and escalation policies?

Creating incident runbooks and escalation policies is handled by generating detailed incident response plans tailored to your service. The Skill produces playbooks with escalation procedures and drill guidelines to ensure rapid, coordinated response during production outages.

How do I design chaos engineering experiments to test system resilience?

Designing chaos engineering experiments involves creating controlled tests that validate steady-state behavior and failure handling. The Skill guides you through resilience testing by modeling experiments to inject failures and verify your system's capacity to maintain production reliability.

Can I use this for capacity planning and auto-scaling configuration validation?

Yes, you can use this for capacity planning and auto-scaling validation. The Skill models load scenarios, plans required infrastructure capacity, and validates your auto-scaling configurations to ensure systems scale correctly under varying demand.

What production context do I need to provide to generate a complete SRE setup?

To generate a complete SRE setup, you simply provide your service context. The Skill then produces a full reliability framework including readiness checks, SLO definitions, runbooks, and incident procedures tailored to your specific production environment.

When do I need SRE governance and readiness reviews across the service lifecycle?

You need SRE governance and readiness reviews when managing systems requiring high availability across their lifecycle. The Skill guides phase-driven readiness checks, monitoring setup, and incident management to maintain resilience from deployment through operations.