observability-and-sre

Coordinate production readiness and ownership across telemetry, alerts, dashboards, and runbooks.

Updated Apr 25, 2026
One-click install
npx skills add https://github.com/Tiepbm/software-engineering-agent --skill observability-and-sre
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observability-and-sre
Source: https://github.com/Tiepbm/software-engineering-agent/tree/main/skills/observability-and-sre
Command: npx skills add https://github.com/Tiepbm/software-engineering-agent --skill observability-and-sre

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Coordinates production supportability, ownership, runbooks, game days, and business workflow visibility while delegating telemetry, alerting, and resilience detail to specialist skills.

Core Features & Use Cases

  • Orchestrates ownership across telemetry, alerts, dashboards, runbooks, and on-call
  • Acts as a launch gate for operability readiness
  • Post-incident review to translate findings into backlog items for specialist skills

Quick Start

Define the service scope, assign owners for telemetry and runbooks, and delegate detailed instrumentation to the appropriate specialist skills.

Frequently Asked Questions about observability-and-sre

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I coordinate production readiness across telemetry, alerts, and runbooks?

You coordinate production readiness by defining service scope, assigning owners for telemetry and runbooks, and delegating detailed instrumentation to specialist skills. This orchestrates ownership across alerts, dashboards, and on-call governance.

When do I need to orchestrate operability readiness for a service launch?

You need operability readiness during design reviews and pre-launch readiness gates. It acts as a launch gate to ensure production supportability, ownership, and business workflow visibility are met before deployment.

What is the best way to translate post-incident review findings into actionable backlog items?

The best way to handle post-incident reviews is to translate findings into backlog items delegated to specialist skills. This ensures telemetry, alerting, and resilience details are addressed by the appropriate experts.

Can I use this for on-call governance without manually setting up every alert and dashboard?

Yes, you can use this for on-call governance because it delegates instrumenting, alerting, and resilience details to appropriate specialist skills. It orchestrates ownership without requiring you to manually build every component.

Does production readiness coordination cover game days and business workflow visibility?

Yes, production readiness coordination covers game days and business workflow visibility. It orchestrates these alongside runbooks and ownership while delegating telemetry and resilience specifics to specialist skills.