Site Reliability Engineer

Automate monitoring, incident response, and post-incident learning for production services.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/tannergolden/repository --skill site-reliability-engineer-tannergolden
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: Site Reliability Engineer
Source: https://github.com/tannergolden/repository/tree/main/.agent/storage/skills/site-reliability-engineering
Command: npx skills add https://github.com/tannergolden/repository --skill site-reliability-engineer-tannergolden

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Systems in production fail and degrade without a consistent reliability and observability strategy. This Skill helps teams design and operate reliable services with proactive monitoring, incident response, and post-incident learning.

Core Features & Use Cases

  • Observability-Driven Reliability: implement end-to-end monitoring, logging, and tracing to detect and prevent outages.
  • Incident Response & Recovery: standardize runbooks, on-call workflows, and rapid recovery across services.
  • SLO/SLA Driven Operation: define and track SLIs and SLOs to meet defined uptime targets and ensure blameless postmortems.
  • Use Case: For a multi-service application, use this Skill to design a reliable deployment pipeline and automated incident playbooks.

Quick Start

Use the site reliability skill to set up a basic reliability monitoring namespace and an initial incident response playbook.

Frequently Asked Questions about Site Reliability Engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I improve system uptime and reliability for multi-service production applications?

Improve system uptime by automating monitoring, incident response, and post-incident learning across multi-service production environments. This approach enforces SLAs, defines SLOs, and maintains robust runbooks to ensure continuous observability and rapid recovery during outages.

What is the best way to set up incident response runbooks for on-call rotations?

The best way to set up incident response runbooks is to standardize on-call workflows and automated recovery playbooks across your services. This ensures rapid incident resolution and consistent recovery processes for every on-call rotation.

How do you define and track SLOs and SLAs to meet uptime targets?

Define and track SLOs and SLAs by establishing clear SLIs and enforcing uptime targets across production services. This framework allows teams to measure reliability metrics, maintain compliance, and drive operational decisions based on actual system performance.

How does blameless postmortem learning work after a production incident?

Blameless postmortem learning works by standardizing post-incident reviews to focus on systemic root causes rather than individual errors. This process captures operational insights from incidents to prevent future outages and improve overall system reliability.

Can I use observability-driven monitoring to detect and prevent production outages?

Yes, implement end-to-end observability-driven monitoring, logging, and tracing to detect and prevent production outages. This proactive strategy provides the visibility needed to identify system degradation early and maintain continuous service availability.

What do I need to design a reliable deployment pipeline with automated incident playbooks?

Designing a reliable deployment pipeline requires defining strict SLOs, establishing continuous observability, and creating standardized incident playbooks. This setup enables automated monitoring and rapid recovery across all deployed production services.