operating-production-services

Define SLOs, write postmortems, and implement on-call practices for production services.

Updated Jan 6, 2026
One-click install
npx skills add https://github.com/salmanparacha/speckitplus-calculator --skill operating-production-services-salmanparacha
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: operating-production-services
Source: https://github.com/salmanparacha/speckitplus-calculator/tree/main/.claude/skills-nocontext/operating-production-services
Command: npx skills add https://github.com/salmanparacha/speckitplus-calculator --skill operating-production-services-salmanparacha

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

SRE patterns for production service reliability: SLOs, error budgets, postmortems, and incident response to help teams define targets and respond effectively.

Core Features & Use Cases

  • Define and manage SLOs and error budgets to measure reliability.
  • Write blameless postmortems and run incident response playbooks.
  • Establish on-call practices and runbooks for rapid recovery across services.

Quick Start

Apply the standard SRE templates to your service today to bootstrap SLOs, postmortems, and incident response.

Frequently Asked Questions about operating-production-services

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLOs and error budgets for my production services?

Define SLOs and error budgets using standard SRE templates to measure service reliability. Apply these patterns to establish clear reliability targets and track error budgets across your services.

What is the best way to write a blameless postmortem after an incident?

Write a blameless postmortem using incident response templates that focus on systemic causes rather than individual fault. Apply SRE playbooks to document timelines, impacts, and action items effectively.

How do I set up on-call practices and runbooks for rapid recovery?

Establish on-call practices and runbooks by applying SRE templates to your services. This streamlines incident response procedures and ensures rapid recovery during production outages.

Can I use these SRE templates for burn-rate alerting across multiple services?

Yes, you can implement burn-rate alerting using the provided SRE templates. The templates satisfy requirements for SRE governance and clear metrics, enabling consistent alerting across multiple services.

Do I need prior SRE experience to bootstrap incident management playbooks?

No prior SRE experience is required to bootstrap incident management. You can apply the standard SRE templates to quickly set up SLOs, postmortems, and incident response practices for your service today.