site-reliability-engineer

Configure Prometheus and Grafana monitoring stacks and manage SLI/SLO definitions.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/imudak/claudemd-gen --skill site-reliability-engineer-imudak
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: site-reliability-engineer
Source: https://github.com/imudak/claudemd-gen/tree/main/.claude/skills/site-reliability-engineer
Command: npx skills add https://github.com/imudak/claudemd-gen --skill site-reliability-engineer-imudak

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill addresses the critical need for robust production monitoring, observability, and incident response to maintain high availability and performance of systems.

Core Features & Use Cases

  • Observability Setup: Configure monitoring tools (Prometheus, Grafana, Datadog) and implement logging, metrics, and tracing.
  • SLO/SLI Management: Define, track, and manage Service Level Objectives and Indicators to ensure service quality.
  • Incident Response: Develop and execute incident response plans, including runbooks and post-mortems.
  • Use Case: When a critical service experiences unexpected downtime, this Skill can help diagnose the issue using observability data, guide the response team through mitigation steps via runbooks, and facilitate a blameless post-mortem to prevent recurrence.

Quick Start

Use the site-reliability-engineer skill to set up Prometheus and Grafana for monitoring.

Frequently Asked Questions about site-reliability-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I set up Prometheus and Grafana for production monitoring?

Set up Prometheus and Grafana to establish production monitoring by configuring metrics collection, logging, and tracing. This provides full observability into system health and visualizes service reliability data.

How do I define and track SLI and SLO for my services?

Define and track SLI and SLO to ensure service quality by establishing clear Service Level Indicators and Objectives. This allows you to measure reliability metrics against targets and maintain uptime standards.

What is the best way to design an incident response workflow with runbooks?

Design an incident response workflow with runbooks to guide mitigation steps during unexpected downtime. This structured approach uses observability data for diagnosis and facilitates blameless post-mortems to prevent recurrence.

Can I use this to manage observability across Datadog, Prometheus, and Grafana?

Yes, you can manage observability across Datadog, Prometheus, and Grafana. The skill configures these monitoring stacks to implement logging, metrics, and tracing for comprehensive production reliability.

Why do I need SLO management for maintaining system uptime?

You need SLO management for maintaining system uptime because it establishes measurable reliability targets through Service Level Objectives. Tracking these indicators proactively addresses system health and prevents service degradation.