observability-design

Design OpenTelemetry-first observability plans with instrumentation, alerts, and retention guidance.

Updated Mar 29, 2026
One-click install
npx skills add https://github.com/jamesogunsan/prod-eng-skills --skill observability-design-jamesogunsan
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-design
Source: https://github.com/jamesogunsan/prod-eng-skills/tree/main/plugins/observability/skills/observability-design
Command: npx skills add https://github.com/jamesogunsan/prod-eng-skills --skill observability-design-jamesogunsan

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Design and improve the observability posture of modern services by aligning telemetry, instrumentation, alerting, and SLI/SLO thinking to enable fast detection and diagnosis.

Core Features & Use Cases

  • Telemetry strategy guidance across metrics, logs, traces, and events
  • OpenTelemetry-first instrumentation recommendations and standardized trace schemas
  • Alerting design, burn-rate policies, and runbooks for operator workflows
  • Use Case: refactor a service's monitoring to reduce alert fatigue and improve issue resolution times

Quick Start

Generate a phased observability design plan focusing on instrumentation, dashboards, and alerting for the target service.

Frequently Asked Questions about observability-design

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design an end-to-end observability plan for my services?

An end-to-end observability plan identifies telemetry gaps and applies an OpenTelemetry-first strategy to review dashboards, traces, logs, and alerts. It defines concrete instrumentation, alert thresholds, retention guidance, and ownership mapping across services.

What is the best way to reduce alert fatigue and improve issue resolution times?

To reduce alert fatigue and improve resolution times, refactor your service monitoring by designing burn-rate alerting policies and runbooks for operator workflows. This aligns SLI/SLO thinking with telemetry to enable fast detection and diagnosis.

How do I identify gaps in my service telemetry and instrumentation?

Identifying telemetry gaps involves reviewing existing metrics, logs, traces, and events against service requirements. The process applies an OpenTelemetry-first instrumentation strategy to standardize trace schemas and ensure coverage across services.

Can I use OpenTelemetry to standardize trace schemas and instrumentation across services?

Yes, OpenTelemetry can standardize trace schemas and instrumentation across services. An OpenTelemetry-first observability plan provides concrete instrumentation recommendations to align telemetry, alerting, and SLI/SLO thinking for fast detection and diagnosis.

What should be included in an observability rollout plan for modern services?

An observability rollout plan should include phased instrumentation steps, dashboard designs, alert thresholds, burn-rate policies, retention guidance, and ownership mapping. It aligns telemetry and operator workflows to improve the observability posture of modern services.

When do I need to refactor a service's monitoring and observability posture?

You need to refactor a service's monitoring when facing alert fatigue, slow issue resolution, or misaligned SLI/SLO tracking. Redesigning observability aligns telemetry, instrumentation, and alerting to enable fast detection and diagnosis across services.