sre-engineer

Defines and tracks SLOs, configures monitoring, and automates operational tasks.

Updated May 31, 2026
One-click install
npx skills add https://github.com/fanguyun/SkillManager --skill sre-engineer-fanguyun
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre-engineer
Source: https://github.com/fanguyun/SkillManager/tree/main/sre-engineer
Command: npx skills add https://github.com/fanguyun/SkillManager --skill sre-engineer-fanguyun

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires prometheus, kubernetes, python, go, terraform, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill unit provides a comprehensive set of tools and best practices for Site Reliability Engineers (SREs) to manage production systems more effectively, reduce toil, and maintain high service reliability.

Core Features & Use Cases

  • Service Level Objective (SLO) Management: Define and track SLOs for availability, latency, and error budgets.
  • Error Budget Policies: Calculate and manage error budgets to plan for and respond to system failures.
  • Monitoring and Alerting: Implement golden signals monitoring and configure alerting based on SLOs.
  • Automation: Automate repetitive tasks and toil reduction with scripts and tools.
  • Chaos Engineering: Design and execute chaos experiments to test system resilience and recovery.
  • Incident Response: Develop incident response procedures, runbooks, and postmortems for improved MTTR.

Quick Start

Load the sre-engineer skill to start defining SLOs, creating error budget policies, and automating incident response procedures.

Frequently Asked Questions about sre-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
Can I configure Prometheus monitoring and alerting based on SLOs?

Yes, you can configure Prometheus monitoring and alerting based on SLOs. It implements golden signals monitoring to track system health and triggers alerts when your service levels are breached.

How do I design and execute chaos engineering experiments to test system resilience?

The best way to reduce operational toil is by automating repetitive tasks with Python and Go scripts. This approach minimizes manual intervention and streamlines incident response procedures.

Do I need to know Terraform and Kubernetes before using this SRE approach?

You design and execute chaos engineering experiments to actively test system resilience and recovery. This validates your infrastructure's ability to withstand failures under controlled conditions.

How do I create incident response runbooks to improve MTTR?

Yes, you need familiarity with Kubernetes, Terraform, Prometheus, and SRE best practices. This foundational knowledge is required to effectively manage production system reliability and define infrastructure policies.

How do I create incident response runbooks to improve MTTR?

You develop incident response procedures, runbooks, and postmortems to improve MTTR. This structured documentation guides responders through standardized recovery steps during production system failures.