observability-sre

Define observability signals, alerts, health checks, and rollback plans for production services.

22|6|Updated Mar 10, 2026
One-click install
npx skills add https://github.com/felvieira/claude-skills-fv --skill observability-sre-felvieira
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-sre
Source: https://github.com/felvieira/claude-skills-fv/tree/main/skills/20-observability-sre
Command: npx skills add https://github.com/felvieira/claude-skills-fv --skill observability-sre-felvieira

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill helps teams design and implement effective observability and reliability practices for production services. It guides you to define structured logs, metrics, tracing, alerts, health checks, readiness probes, error budgets, rollback strategies, and safe operational procedures.

Core Features & Use Cases

  • Define essential signals (logs, metrics, traces) and actionable alerts.
  • Establish health checks, readiness probes, rollback plans, and incident response runbooks.
  • Align across deployment pipelines to improve reliability and operability with governance and guardrails.

Quick Start

Draft an observability plan that covers logs, metrics, tracing, and rollback steps for your service.

Frequently Asked Questions about observability-sre

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I set up observability for production services with logs, metrics, and tracing?

To implement observability, define structured logs, metrics, and traces to capture system behavior. You must then establish actionable alerts based on these signals to detect production issues early.

What is the best way to design incident response runbooks and rollback strategies?

Designing incident response runbooks involves creating safe operational procedures and rollback plans. This Skill helps you establish rollback capabilities across deployment pipelines and define governance guardrails for safe operations.

How do I configure health checks and readiness probes for reliable service deployment?

Configuring health checks and readiness probes ensures services are operational before receiving traffic. This Skill provides guidance on establishing these checks alongside deployment pipelines to improve operability and reliability.

Does this observability guidance work for systems requiring error budgets and governance?

Yes, this observability guidance supports systems requiring error budgets and governance. It helps teams align across deployment pipelines with operational guardrails to ensure safe, reliable production operations.

How to create an observability plan covering tracing, metrics, and rollback steps?

To create an observability plan, draft a comprehensive strategy covering logs, metrics, tracing, and rollback steps for your service. This ensures structured signals and incident handling procedures are properly established.