production-monitoring

Assess production health and configure monitoring and alerting for deployed services.

Updated Feb 28, 2026
One-click install
npx skills add https://github.com/dsivov/ai_development_team --skill production-monitoring-dsivov
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: production-monitoring
Source: https://github.com/dsivov/ai_development_team/tree/main/.claude/skills/production-monitoring
Command: npx skills add https://github.com/dsivov/ai_development_team --skill production-monitoring-dsivov

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Production environments require continuous visibility, reliable health checks, and timely alerts to prevent outages. This Skill provides a structured approach to monitoring readiness, performance, and incident response across services.

Core Features & Use Cases

  • Health endpoint verification and readiness checks
  • Definition of golden signals and SLIs with alerting guidelines
  • Post-deployment verification and incident documentation templates

Quick Start

Configure monitoring for your service and verify endpoints.

Frequently Asked Questions about production-monitoring

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I set up health checks and SLA tracking for cloud-native services?

To set up health checks and SLA tracking, you verify readiness endpoints and define golden signals as SLIs. This provides continuous production visibility and structured alerting guidelines to maintain service reliability.

What's the best way to establish alerting guidelines for production incidents?

The best way to establish alerting guidelines is by defining golden signals and SLIs specific to your deployed services. This ensures timely, actionable alerts and provides templates for post-incident documentation.

Can I use this approach for post-deployment verification and release validation?

Yes, you can use this approach for post-deployment verification and release validation. It applies health endpoint verification to confirm service readiness and operational stability across cloud-native architectures.

What are golden signals and how do they apply to observability?

Golden signals are core metrics like latency, traffic, errors, and saturation used for observability. They define SLIs that establish alerting guidelines, ensuring continuous visibility into production health and incident response.

Does this monitoring framework require specific dependencies for cloud-native architectures?

No specific dependencies are required to apply this monitoring framework. It operates independently to establish consistent observability, SLA tracking, and incident response across various cloud-native architectures.

How do I document production incidents after they occur?

You document production incidents using provided post-incident documentation templates. This standardizes the recording of incident response actions, health endpoint failures, and SLA breaches within the operational framework.