ops-observability-slo

Define SLO metrics, alert thresholds, and incident review processes.

3|Updated Mar 11, 2026
One-click install
npx skills add https://github.com/junchenghuo/openclaw-biz-agent --skill ops-observability-slo
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ops-observability-slo
Source: https://github.com/junchenghuo/openclaw-biz-agent/tree/main/ops/.agents/skills/ops-observability-slo
Command: npx skills add https://github.com/junchenghuo/openclaw-biz-agent --skill ops-observability-slo

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps establish a robust observability system and define quantifiable service quality metrics, ensuring services meet defined performance and reliability standards.

Core Features & Use Cases

  • Metric Definition: Define key performance indicators (KPIs) for availability, latency, and error rates.
  • Alerting Configuration: Set up tiered alerting rules and suppression mechanisms.
  • Observability Integration: Link logs, traces, and metrics for comprehensive monitoring.
  • Reporting: Generate weekly reports and incident review templates.
  • Use Case: A team can use this Skill to define the Service Level Objectives (SLOs) for their new microservice, configure alerts for when these SLOs are at risk, and establish a process for reviewing any incidents that occur.

Quick Start

Use the ops-observability-slo skill to define SLOs for service availability and latency.

Frequently Asked Questions about ops-observability-slo

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLOs for microservice availability and latency?

To define SLOs for microservice availability and latency, establish key performance indicators for error rates and response times, then configure alert thresholds to monitor when these service level objectives are at risk.

What is observability and how does log-trace-metric correlation improve incident management?

Observability links logs, traces, and metrics to provide comprehensive monitoring of system health. This correlation improves incident management by enabling operational teams to quickly identify performance bottlenecks and review incidents.

How do I set up tiered alerting rules and suppression mechanisms for performance metrics?

Set up tiered alerting rules and suppression mechanisms by defining alert policies based on your service level objectives. This ensures operational teams receive appropriately prioritized notifications for availability and error rate threshold breaches.

Can I generate weekly reports and incident review templates from existing performance metrics?

Yes, you can generate weekly reports and incident review templates from existing performance metrics. This reporting capability helps operational teams document system health, review incidents, and ensure services meet reliability standards.

What's the best way to establish a quantifiable service quality monitoring process for a new service?

The best way to establish quantifiable service quality monitoring is to define availability, latency, and error rate indicators with associated alert policies. This creates a robust observability system ensuring services meet performance standards.