observability-slo-design

Design service-level objectives, error budgets, burn-rate alerts, dashboards, and runbooks.

1|Updated Jul 17, 2026
One-click install
npx skills add https://github.com/Arafly/sre-playbooks --skill observability-slo-design
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-slo-design
Source: https://github.com/Arafly/sre-playbooks/tree/main/observability-slo-design
Command: npx skills add https://github.com/Arafly/sre-playbooks --skill observability-slo-design

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps you move from machine-focused monitoring to user-focused reliability by defining what healthy really means for a service. It replaces noisy, low-signal alerts with meaningful SLIs, SLOs, error budgets, dashboards, and runbooks tied to customer impact.

Core Features & Use Cases

  • Monitoring inventory and gap analysis: Review existing dashboards, alerts, runbooks, and incident pain points to find missing signals and vanity metrics.
  • Critical journey and SLO design: Identify the most important user journeys, choose the right SLIs, set targets, and document error-budget policy and ownership.
  • Burn-rate alerting and operational readiness: Design fast-burn and slow-burn alerts, validate dashboards and runbooks, and ensure on-call responders can act from a fresh context.

Quick Start

Help me design or review the SLOs, alerts, dashboards, and runbooks for this service so we can focus on user-impacting reliability rather than noisy infrastructure metrics.

Frequently Asked Questions about observability-slo-design

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I replace noisy infrastructure alerts with user-impacting SLOs?

To replace noisy infrastructure alerts with user-impacting SLOs, define service-level indicators tied to critical user journeys and establish error budgets. This shifts focus from machine-level dashboards to actual customer impact, reducing alert fatigue and improving on-call actionability.

What is an error budget and how does it improve incident response?

An error budget is the allowable threshold of unreliability for a service based on your SLO targets. It improves incident response by providing a clear burn-rate metric, allowing on-call responders to prioritize alerts based on actual user impact rather than noisy vanity metrics.

How do I design burn-rate alerts for my critical user journeys?

Design burn-rate alerts by setting fast-burn thresholds for rapid error spikes and slow-burn thresholds for sustained issues against your SLO error budgets. Validate these alerts with runbooks so on-call responders can act decisively from a fresh context during an incident.

When do I need to define SLIs and SLOs for my service?

You need to define SLIs and SLOs when your service suffers from noisy monitoring, missing reliability targets, or unclear critical journeys. This process establishes baseline measurements and ownership clarity required to transition from vanity metrics to actionable reliability signals.

What is the best way to inventory existing monitoring and find missing reliability signals?

The best way to inventory existing monitoring is to review current dashboards, alerts, and runbooks to identify gap analysis and incident pain points. This helps find missing signals, eliminate vanity metrics, and prioritize critical journeys for SLO design.

Why are my current dashboards generating low-signal alerts and causing alert fatigue?

Dashboards generate low-signal alerts and cause alert fatigue when they rely on machine-focused infrastructure metrics instead of user-focused reliability. Defining SLIs and SLOs tied to customer impact replaces noisy alerts with meaningful burn-rate alerts and actionable runbooks.