observability-designer

Design SLI/SLO frameworks, alert rules, and dashboards for software systems.

2|Updated Mar 13, 2026
One-click install
npx skills add https://github.com/zhangzhang-111-i/claude-skills111 --skill observability-designer-zhangzhang-111-i
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-designer
Source: https://github.com/zhangzhang-111-i/claude-skills111/tree/main/engineering/observability-designer
Command: npx skills add https://github.com/zhangzhang-111-i/claude-skills111 --skill observability-designer-zhangzhang-111-i

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill addresses the challenge of creating comprehensive and effective observability strategies, ensuring systems are monitored reliably and efficiently.

Core Features & Use Cases

  • SLO Framework Design: Automatically generate Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Alert Optimization: Analyze and refine existing alerts to reduce noise and improve actionability.
  • Dashboard Generation: Create role-specific dashboards for SREs, developers, and executives.
  • Use Case: A new microservice is launched. Use this Skill to automatically generate its SLOs, design essential monitoring dashboards, and create optimized alert rules, ensuring it's production-ready from day one.

Quick Start

Use the observability-designer skill to generate an SLO framework for a new critical API service.

Frequently Asked Questions about observability-designer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design SLOs and SLIs for a new microservice?

To design SLOs for a microservice, you define Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to set reliability targets. This process automatically generates error budgets and tailored metrics for APIs, databases, and queues.

What is the best way to reduce alert noise and improve actionability?

The best way to reduce alert noise is through alert optimization, which analyzes and refines existing alert rules. This ensures alerts are highly actionable and tailored to specific service types like APIs, web applications, and queues.

Can I generate monitoring dashboards for different roles like SREs and developers?

Yes, you can generate monitoring dashboards tailored for specific roles such as SREs, developers, and executives. This ensures each persona views the most relevant metrics, logs, and traces for their operational context.

Does this observability strategy approach work for database and queue services?

Yes, this observability strategy specifically supports database and queue services alongside API and web services. It provides tailored recommendations for metrics, logging, and tracing to ensure comprehensive monitoring across these distinct architectures.

How do I make a new API service production-ready with monitoring from day one?

To make a new API production-ready, you generate a complete observability strategy that includes SLO frameworks, essential monitoring dashboards, and optimized alert rules. This ensures reliable and efficient system monitoring from launch.

When do I need to define an error budget for my software system?

You need to define an error budget when establishing Service Level Objectives (SLOs) for your software system. It quantifies the allowable unreliability, helping teams balance feature velocity with system stability based on tailored metrics.