sre-operations-lead

Establish SRE alerting and incident response practices for Prometheus/Grafana/Loki/Tempo stacks.

8|1|Updated Mar 30, 2026
One-click install
npx skills add https://github.com/drewid74/ai_skills --skill sre-operations-lead
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre-operations-lead
Source: https://github.com/drewid74/ai_skills/tree/main/sre-operations-lead
Command: npx skills add https://github.com/drewid74/ai_skills --skill sre-operations-lead

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps you stop alert fatigue and restore reliability by designing actionable observability, incident response, and capacity planning practices.

Core Features & Use Cases

  • Observability stack strategy: Choose metrics, logs, and traces to answer real reliability questions like “why is it slow?”
  • Alert design and SLO-backed governance: Define SLIs/SLOs and craft alerts that correlate with user impact, include runbooks, and reduce flapping.
  • Incident operations and continuous improvement: Use an incident workflow (triage → mitigate → investigate → postmortem) and enforce action items.

Quick Start

Ask an AI to set up SLOs, Prometheus alert rules with runbook annotations, Grafana dashboards provisioned from git, and an incident postmortem template for your service.

Frequently Asked Questions about sre-operations-lead

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I reduce Prometheus alert fatigue and stop noisy flapping alerts?

Reduce Prometheus alert fatigue by defining SLO-backed alert rules that correlate directly with user impact. Enforce runbook-first annotations and configure Alertmanager routing and grouping to suppress flapping notifications, ensuring only actionable alerts reach responders.

How do I define SLIs and SLOs for incident response workflows?

Define SLIs and SLOs using PromQL-based metrics to measure service reliability against user impact. Apply these SLOs to incident response workflows spanning triage, mitigation, investigation, and postmortem stages, enforcing continuous improvement through action items.

What's the best way to structure Grafana dashboards and observability stacks for reliability?

Structure observability stacks across Prometheus, Grafana, Loki, and Tempo to answer reliability questions like why a service is slow. Provision Grafana dashboards from git and apply tail sampling strategies for traces to ensure observability data remains actionable.

Does this SRE operations approach work with existing Alertmanager routing and capacity planning?

Yes, this SRE operations approach integrates with existing Alertmanager routing and capacity planning. It enforces quality gates for alert design and applies capacity forecasting within Prometheus and Grafana-based monitoring stacks to improve overall reliability.

Why do my current monitoring alerts lack runbook annotations and postmortem action items?

Monitoring alerts lack runbook annotations and postmortem action items when SRE quality gates are not enforced. Establish runbook existence requirements for alert rules and track postmortem action items to completion to transition from noisy alerts to calm operations.

Can I use this to set up incident postmortem templates and capacity forecasting for my service?

Yes, you can generate incident postmortem templates and apply capacity forecasting for your service. It establishes SRE practices that define observability strategy and capacity planning to reduce alert noise while improving long-term reliability.