observability-planner

Plan and implement production observability with metrics, alerts, dashboards, and SLOs.

50|9|Updated Feb 5, 2026
One-click install
npx skills add https://github.com/Agile-V/agile_v_skills --skill observability-planner
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-planner
Source: https://github.com/Agile-V/agile_v_skills/tree/main/observability-planner
Command: npx skills add https://github.com/Agile-V/agile_v_skills --skill observability-planner

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Production systems require measurable visibility to detect, diagnose, and prevent incidents; this Skill defines a structured observability plan to tie metrics, events, dashboards, alerts, and SLOs to requirements.

Core Features & Use Cases

  • Define MET-XXXX metrics, ALR-XXXX alerts, SLO-XXXX objectives, and incident-adjacent governance.
  • Build dashboards that surface per-REQ validation and system health.
  • Enable incident feedback loops to drive continuous improvement across cycles.

Quick Start

Create an initial OBSERVABILITY_PLAN.md outlining MET-XXXX metrics, ALR-XXXX alerts, and SLO-XXXX objectives for your production service.

Frequently Asked Questions about observability-planner

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design production observability metrics and alerts for cloud-native systems?

To design production observability, you plan structured metrics, events, dashboards, alerts, and SLOs tied directly to requirements. This enables per-requirement validation and incident-driven improvement for cloud-native systems.

What is the best way to define SLOs and alerts that align with production requirements?

Defining SLOs and alerts requires mapping governance objectives like SLO-XXXX and ALR-XXXX to specific requirements. This structured approach ensures alerts trigger incident feedback loops for continuous production improvement.

How do I create a dashboard that surfaces per-requirement validation and system health?

You create dashboards by building reference outputs that visualize both overall system health and per-requirement validation. These dashboards provide measurable visibility to detect and diagnose production incidents.

When do I need to implement an observability plan for production operations?

You need an observability plan during release activities at Gate 2 and beyond, as well as for continuous production operations. It establishes measurable visibility to detect, diagnose, and prevent incidents.

Can I use this approach to establish incident feedback loops for continuous improvement?

Yes, you can establish incident feedback workflows by defining governance around incidents and producing runbooks. This enables incident-driven improvement across cycles to continuously refine observability.

How do I start planning observability if I only have basic production metrics?

You start by creating an OBSERVABILITY_PLAN.md file outlining your MET-XXXX metrics, ALR-XXXX alerts, and SLO-XXXX objectives. This initial plan provides the structured foundation for your production service.