sre-engineer

Guide production incident triage and remediation with SLO and observability practices.

Updated Mar 30, 2026
One-click install
npx skills add https://github.com/Patkik/Multi-tenant-SaaS-Catering-V2 --skill sre-engineer-patkik
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre-engineer
Source: https://github.com/Patkik/Multi-tenant-SaaS-Catering-V2/tree/main/.agents/skills/sre-engineer
Command: npx skills add https://github.com/Patkik/Multi-tenant-SaaS-Catering-V2 --skill sre-engineer-patkik

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill helps teams quickly detect, triage, and mitigate production incidents, design and enforce SLOs/SLAs, and harden reliability across systems.

Core Features & Use Cases

  • Incident triage and rollback-safe mitigation to reduce blast radius during outages.
  • SLOs, error budgets, and alert quality improvements to align reliability with business goals.
  • Logging, metrics, tracing, and dashboards to improve observability and on-call efficiency.
  • Deployment-time guardrails and runbook automation to prevent regressions and speed recovery.

Quick Start

Describe a current incident or reliability goal to start guided triage and remediation.

Frequently Asked Questions about sre-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I triage production incidents and reduce blast radius during outages?

Triage production incidents through guided mitigation and rollback-safe remediation to reduce blast radius during outages. Describe a current incident to start structured guidance and quickly restore system stability.

How do I design SLOs and error budgets to align reliability with business goals?

Design SLOs and error budgets by applying reliability guardrails to align availability targets with business goals. This skill enforces alert quality improvements and tracks error budgets to prevent regression during deployments.

What is the best way to automate runbooks and add deployment guardrails for reliability?

Automate runbooks and add deployment guardrails by applying reliability checks during deploys and runtime. This prevents regressions and speeds recovery through structured incident response and automated mitigation steps.

How can I improve observability and on-call efficiency using metrics, tracing, and logging?

Improve observability and on-call efficiency by implementing logging, metrics, tracing, and dashboards across production systems. This provides actionable visibility and reduces mean time to resolution during incidents.

Can I use this guided incident response approach for any production system scale?

You can use this guided incident response approach for production systems at scale. It hardens reliability across architectures through SLO enforcement, observability improvements, and rollback-safe mitigation without platform restrictions.

When should I not use automated runbook automation for incident response?

You should not use automated runbook automation for incident response when a production issue requires novel, unscripted debugging outside existing guardrails. It works best for structured, predictable mitigation paths and SLO enforcement.