sre-patterns

Document SRE patterns for SLOs, SLIs, runbooks, and incident response.

69|9|Updated Jan 30, 2026
One-click install
npx skills add https://github.com/Tibsfox/gsd-skill-creator --skill sre-patterns
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre-patterns
Source: https://github.com/Tibsfox/gsd-skill-creator/tree/main/examples/skills/sre-patterns
Command: npx skills add https://github.com/Tibsfox/gsd-skill-creator --skill sre-patterns

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

SRE-patterns provide a structured approach to building reliable software systems by codifying objectives, measurement, and runbooks to prevent outages and toil.

Core Features & Use Cases

  • Establish SLOs/SLIs/SLA definitions and measurement frameworks to align engineering and business expectations.
  • Reduce toil through automation, runbooks, and capacity planning with scalable patterns.
  • Guide incident response, postmortems, and reliability reviews to improve resilience across services.

Quick Start

Define a basic SLO for a service, implement a simple availability alert, and document a runbook for incident response.

Frequently Asked Questions about sre-patterns

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLOs and SLIs for my service?

Defining SLOs and SLIs requires establishing measurement frameworks that align engineering and business expectations. This approach codifies service level objectives and indicators to standardize reliability tracking across engineering teams.

What is the best way to reduce operational toil in engineering teams?

Reducing operational toil involves applying automation, runbooks, and capacity planning with scalable SRE patterns. This structured approach prevents repetitive manual work by codifying incident response and reliability reviews.

How do I conduct a postmortem after an incident?

Conducting a postmortem is guided by established SRE patterns that document incident response and reliability reviews. This process improves resilience across services by systematically analyzing outages and applying governance checks.

When do I need capacity planning for my software systems?

Capacity planning is needed when building reliable software systems that require scalable SRE patterns to prevent outages. It helps organizations manage resources proactively and align engineering capacity with service level objectives.

Can I use these SRE patterns for governance checks across multiple teams?

Yes, these SRE patterns satisfy requirements for documenting SRE concepts and performing governance checks across engineering teams. They provide a structured approach to standardize reliability and align business expectations organization-wide.