observe-production

Evaluate deployed service health using SLOs, error rates, latency, throughput, and alerts.

5|Updated Jul 25, 2025
One-click install
npx skills add https://github.com/tomzx/agents --skill observe-production
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: observe-production
Source: https://github.com/tomzx/agents/tree/main/skills/observe-production
Command: npx skills add https://github.com/tomzx/agents --skill observe-production

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Checks the health of a deployed service or feature by reviewing SLOs, error rates, latency, throughput, and recent alerts. Produces a health report suitable for maintenance reviews, post-deploy verification, or incident triage.

Core Features & Use Cases

  • Monitor SLOs/SLIs and alert status across services.
  • Assess error rates, latency percentiles, throughput, and recent alerts for incident triage.
  • Produce a structured health report for maintenance reviews, post-deploy verification, or incident response.

Quick Start

Check the health of the deployed service and generate a concise production health report.

Frequently Asked Questions about observe-production

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I check the health of a deployed service after a release?

To check deployed service health, evaluate SLOs, error rates, latency, throughput, and recent alerts. This Skill aggregates observability data to produce a structured health report suitable for post-deploy verification and incident triage.

What metrics do I need to assess production service health for incident triage?

Assessing production service health requires reviewing SLOs, error rates, latency percentiles, throughput, and recent alerts. Aggregating these metrics helps generate a structured health report for incident triage and routine maintenance reviews.

Can I use this to monitor SLO and SLI status across individual services?

Yes, you can monitor SLO and SLI status across individual services or features. It evaluates error rates, latency, throughput, and alert status to generate a structured health report for maintenance reviews and incident response.

How do I generate a structured health report from observability data?

Generate a structured health report by aggregating observability data from monitoring tools to evaluate SLOs, error rates, latency, and throughput. This produces a concise report suitable for maintenance reviews and troubleshooting.

What is the best way to verify SLO compliance and detect production issues fast?

The best way to verify SLO compliance and detect production issues is by evaluating service health through error rates, latency, throughput, and alerts. This approach produces a structured health report for fast incident triage and troubleshooting.