observability-designer

Design SLI and SLO frameworks, alert optimization, and dashboard specifications for software services.

Updated Apr 24, 2026
One-click install
npx skills add https://github.com/Veloxia-agency/VELOXIA-WEB --skill observability-designer-veloxia-agency
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-designer
Source: https://github.com/Veloxia-agency/VELOXIA-WEB/tree/main/.claude/skills/engineering/skills/observability-designer
Command: npx skills add https://github.com/Veloxia-agency/VELOXIA-WEB --skill observability-designer-veloxia-agency

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill helps teams turn raw service signals into production-ready observability plans, reducing guesswork around reliability, performance, and incident response.

Core Features & Use Cases

  • SLI and SLO Framework Design: Define meaningful indicators, set targets, and calculate error budgets for services that matter to users.
  • Alert Optimization: Review alert rules for noise, missing coverage, duplicate conditions, and missing runbook context.
  • Dashboard Generation: Produce role-aware monitoring dashboards for SRE, developer, executive, and operations workflows.
  • Use Case: A platform team can use this Skill to onboard a new payment service, create its SLOs, tune alerts, and ship a Grafana-ready dashboard in one workflow.

Quick Start

Ask the observability designer skill to analyze my service and generate an SLI and SLO framework, an alert optimization plan, and a dashboard specification.

Frequently Asked Questions about observability-designer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design SLI and SLO frameworks for production services?

Design SLI and SLO frameworks by defining meaningful service indicators, setting reliability targets, and calculating error budgets. Provide service definitions and monitoring context to generate deterministic, Python-backed recommendations with operational guidance for your applications.

What is the best way to optimize alerting rules and reduce noise in Prometheus?

Optimize alerting by reviewing existing rules for noise, missing coverage, and duplicate conditions. Input your alert configurations to receive an optimization plan that adds missing runbook context and ensures your alerts are actionable.

How do I generate role-aware dashboards for APIs, web apps, and databases?

Generate role-aware dashboards by providing service definitions and monitoring context to produce dashboard specifications tailored for SRE, developer, executive, and operations workflows. This yields Grafana-ready dashboard layouts for APIs, databases, and queues.

Can I use this to build observability strategies for batch jobs and machine learning services?

Yes, you can build observability strategies for batch jobs and machine learning services. The skill designs monitoring plans, defines SLIs, and generates alert configurations specifically tailored for these specialized workloads alongside standard web apps.

Do I need existing monitoring context to create a production observability plan?

Yes, providing service definitions, alert configurations, and monitoring context is required. The skill uses these inputs to produce deterministic recommendations, ensuring the resulting observability strategy matches your actual production environment.

Why do my current Grafana dashboards lack role-aware visibility for my operations team?

Dashboards often lack role-aware visibility because they are not tailored to specific workflows. Generate role-aware dashboard specifications for SRE, developer, executive, and operations teams by supplying your service definitions and monitoring context.