agent-sre

Automate production reliability with observability, incident response, and postmortems.

4|Updated Mar 31, 2026
One-click install
npx skills add https://github.com/ryan-nguyen-01/agent-platform --skill agent-sre
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: agent-sre
Source: https://github.com/ryan-nguyen-01/agent-platform/tree/main/.claude/Agents/agent-sre
Command: npx skills add https://github.com/ryan-nguyen-01/agent-platform --skill agent-sre

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

SRE-empowered agent ensures production reliability by designing monitoring, SLI/SLOs, alerting, incident response, and postmortems to minimize downtime and accelerate recovery.

Core Features & Use Cases

  • Observability setup with metrics, logs, and traces; alerting and dashboards to guard service health.
  • SLI/SLO definition, error budgets, and runbooks to guide incident handling and recovery.
  • Incident response playbooks and postmortem templates to learn from failures and prevent recurrences.
  • Production readiness checks and collaboration with DevOps and coding agents to ensure safe deployments.

Quick Start

Define your service's monitoring, SLIs/SLAs, and alerting rules, then run a simulated incident to validate response.

Frequently Asked Questions about agent-sre

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLI and SLO metrics for my Kubernetes services?

Define SLI and SLO metrics for Kubernetes services by establishing error budgets and runbooks to guide incident handling and validate service health. This approach directly measures service reliability and accelerates incident recovery.

What is the best way to automate incident response and triage for production outages?

Automate incident response and triage for production outages by applying phase-driven observability, alerting rules, and incident response playbooks to minimize downtime. This ensures rapid detection and structured recovery during service disruptions.

How do I set up observability with metrics, logs, and traces for service monitoring?

Set up observability with metrics, logs, and traces by configuring alerting and dashboards to continuously guard service health. This phase-driven monitoring approach validates production readiness and tracks system reliability.

Can I use this SRE approach for production readiness checks before deploying new features?

You can use this SRE approach for production readiness checks by collaborating with DevOps workflows to ensure safe deployments. It validates monitoring, SLI definitions, and alerting rules before releasing new features.

How do I write a postmortem template to learn from production incidents?

Write a postmortem template to learn from production incidents by documenting failure analysis and preventive actions to avoid recurrences. This structured post-incident learning process improves long-term service reliability.

Does this SRE workflow require specific dependencies to manage production reliability?

This SRE workflow requires no specific dependencies to manage production reliability, applying directly across DevOps and operator workflows. It functions independently to automate monitoring, alerting, and post-incident learning.