sre-engineer

Identify reliability risks and propose minimal verifiable fixes with rollback plans.

22|2|Updated Mar 24, 2026
One-click install
npx skills add https://github.com/jshsakura/awesome-opencode-skills --skill sre-engineer-jshsakura
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sre-engineer
Source: https://github.com/jshsakura/awesome-opencode-skills/tree/main/skills/sre-engineer
Command: npx skills add https://github.com/jshsakura/awesome-opencode-skills --skill sre-engineer-jshsakura

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Own site reliability engineering work as production-safety and operability engineering, not checklist completion.

Core Features & Use Cases

  • Focus on establishing robust SLO-aligned reliability improvements, incident runbooks, and safe rollback strategies.
  • Provide measurable telemetry and guardrails to prevent uncontrolled blast radius.
  • Coordinate across control plane, data plane, and dependencies for end-to-end resilience.

Quick Start

Propose the smallest safe reliability improvement with rollback options for the current incident.

Frequently Asked Questions about sre-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I propose the smallest safe reliability improvement during an incident?

To propose safe reliability improvements, identify the current risk and limit analysis to control plane, data plane, and dependency edges, providing evidence-based rationale with rollback options validated against production telemetry.

What is the best way to establish SLO-aligned incident runbooks?

Establishing SLO-aligned incident runbooks involves coordinating across control plane, data plane, and dependencies to ensure end-to-end resilience, while providing measurable telemetry and guardrails to prevent uncontrolled blast radius.

How does risk analysis work for site reliability engineering?

Risk analysis for site reliability engineering works by identifying current reliability risks and proposing the smallest, verifiable fix that preserves security boundaries, ensuring recommendations reference measurable indicators.

Can I validate SLO recommendations against production telemetry?

Yes, you can validate SLO recommendations against production telemetry, ensuring that proposed reliability improvements include rollback or degrade-path plans and evidence-based rationale for safe deployment.

When do I need rollback strategies for dependency edges?

You need rollback strategies for dependency edges when coordinating end-to-end resilience, ensuring that any proposed safe changes include verifiable rollback options to prevent uncontrolled blast radius.

Why focus on small changes for incident management rather than complete overhauls?

Focusing on small changes for incident management ensures verifiable fixes that preserve security boundaries, limiting blast radius and providing measurable telemetry guardrails instead of risking uncontrolled disruptions during reliability improvements.