monitoring-alerting-and-slos

Design production monitoring, alerting, and SLOs with SLIs, dashboards, and runbooks.

Updated Apr 25, 2026
One-click install
npx skills add https://github.com/Tiepbm/software-engineering-agent --skill monitoring-alerting-and-slos
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: monitoring-alerting-and-slos
Source: https://github.com/Tiepbm/software-engineering-agent/tree/main/skills/monitoring-alerting-and-slos
Command: npx skills add https://github.com/Tiepbm/software-engineering-agent --skill monitoring-alerting-and-slos

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Production systems often suffer from gaps in reliable monitoring, alerting, and governance, leading to delayed incident response and missed business signals.

Core Features & Use Cases

  • Define SLIs, SLOs, dashboards, and runbooks aligned to critical business workflows.
  • Create actionable alerts with owners, severity classifications, and remediation steps.
  • Map health signals to business outcomes across APIs, data pipelines, queues, and batch jobs.
  • Provide dashboards that differentiate service health, dependency health, and incident diagnosis.

Quick Start

Define your service's critical workflows, map SLIs to business goals, and draft initial SLO targets.

Frequently Asked Questions about monitoring-alerting-and-slos

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLIs and SLOs for production services?

To define SLIs and SLOs, map your service's critical workflows to business goals and draft initial targets. This process specifies service level indicators, objectives, alert rules, and runbooks to improve system reliability and ensure incident readiness.

What is the best way to set up actionable alerts for incident management?

Actionable alerts require assigning specific owners, severity classifications, and clear remediation steps. By mapping health signals directly to critical business outcomes across APIs and data pipelines, you ensure fast incident response and reduce delayed signal gaps.

Can I use this approach to monitor data pipelines and batch jobs?

Yes, this approach applies to monitoring data pipelines, batch jobs, and queues. It maps health signals to business outcomes across these workflows, ensuring your dashboards differentiate service health, dependency health, and incident diagnosis effectively.

How do I design health dashboards that help with incident diagnosis?

Design health dashboards by differentiating service health from dependency health to aid incident diagnosis. Align dashboards with critical business workflows, specified SLIs, and runbooks to ensure clear owner accountability during production incidents.

Do I need existing telemetry infrastructure to implement these SLOs?

You need to identify your service's critical workflows to map telemetry to business goals. The implementation specifies SLIs, SLOs, and alert rules based on existing telemetry data from APIs, data pipelines, and queues to establish reliability targets.