sre-engineer

Define SLOs, error budgets, and reliability targets for deployments.

3|2|Updated Feb 27, 2026
One-click install
npx skills add https://github.com/grasberg/sofia --skill sre-engineer-grasberg
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre-engineer
Source: https://github.com/grasberg/sofia/tree/main/workspace/skills/sre-engineer
Command: npx skills add https://github.com/grasberg/sofia --skill sre-engineer-grasberg

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Site reliability engineering tasks such as defining SLOs, budgets, and incident response can be error-prone and time-consuming; this skill provides structured guidance and automation patterns to improve reliability and reduce toil.

Core Features & Use Cases

  • Define SLIs, SLOs, and manage error budgets to guide deployments.
  • Design observability stacks (metrics, logs, traces) and create actionable runbooks for blameless post-mortems.
  • Automate toil and establish guardrails to improve incident response and on-call workflows.

Quick Start

Draft your first SLI/SLO definitions and create a starter runbook to begin improving reliability today.

Frequently Asked Questions about sre-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLOs and error budgets to guide complex system deployments?

You define SLOs and error budgets by establishing structured SLIs and reliability targets that enforce deployment guardrails. This approach applies error budget tracking across incident management and observability design to ensure system resilience.

What is the best way to automate toil and improve on-call workflows for incident response?

Automating toil for incident response involves applying automation patterns and establishing guardrails to improve on-call workflows. This reduces error-prone manual tasks and boosts reliability during complex system incidents.

How do I create actionable runbooks for blameless post-mortems?

Creating actionable runbooks for blameless post-mortems requires designing observability stacks that integrate metrics, logs, and traces. These structured runbooks guide incident management and automate toil for complex systems.

Can I use this approach to design observability stacks for complex systems?

Yes, designing observability stacks for complex systems is a core use case. You apply structured SLIs and SLOs across metrics, logs, and traces to create actionable alert strategies and improve incident response.

Why do I need structured SLIs and SLOs for incident management?

Structured SLIs and SLOs are needed for incident management to provide clear reliability targets and error budget tracking. They enforce alert strategies and automation patterns that reduce error-prone manual interventions.

Does implementing SRE automation patterns require specific platform dependencies?

No specific platform dependencies are required to implement SRE automation patterns. The approach provides structured guidance for reliability targets and observability design that applies across complex systems.