technology-reliability-management

Define SLOs and operate incident, change, capacity, and continuity controls for critical systems.

Updated Aug 22, 2026
One-click install
npx skills add https://github.com/fritzgeraldz/Vibe-Managing --skill technology-reliability-management-fritzgeraldz
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: technology-reliability-management
Source: https://github.com/fritzgeraldz/Vibe-Managing/tree/main/skills/technology/technology-reliability-management
Command: npx skills add https://github.com/fritzgeraldz/Vibe-Managing --skill technology-reliability-management-fritzgeraldz

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve? Founders and operators often lack a structured way to keep critical technology systems dependable: service objectives are undefined, incidents recur, changes break production, and recovery is never tested. This Skill turns reliability work into an evidence-backed decision and operating plan tied to business constraints. ## Core Features & Use Cases - Service tiering and SLO definition: Classify services by criticality, set service level objectives, and instrument reliability signals with confidence-scored evidence. - Incident, problem, and change control: Manage incidents and recurring problems, control change failure rate, and test recovery and continuity procedures. - Reliability economics review: Quantify risk-adjusted value, downside loss, and constraint headroom for reliability investments, simulated in the Business Digital Twin. - Use Case: A founder asks how to reduce outages without exceeding cash and risk limits. The agent tiers services, sets SLOs, compares at least three feasible options against the counterfactual, and recommends the highest confidence-weighted plan with monitoring and escalation rules. ## Quick Start Use technology reliability management to define SLOs and an incident and change control plan for our critical systems within our cash and risk limits.

Frequently Asked Questions about technology-reliability-management

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLOs for critical business systems?▼

Tier services by criticality first, then set service level objectives with baselines, targets, and deadlines tied to business outcomes. Instrument reliability signals and attach evidence, confidence, and a decision implication to each step before committing resources.

How to reduce change failure rate and recurring incidents?▼

Control changes through a structured change management process, manage incidents and root-cause problems systematically, and test recovery procedures regularly. Track change failure rate, recovery time, and recurring incidents as KPIs with owners and thresholds.

When should I not use a reliability management skill?▼

Do not use it during an active emergency before invoking the incident or crisis workflow, and do not use generic benchmarks before comparability is calibrated. Legal, safety, or regulated determinations require a licensed specialist instead.

What context is required before running a reliability review?▼

Load company, goals, strategy, relevant metrics, prior decisions, and the current Digital Twin. Run archetype, industry profile, stage and maturity, and regulatory intensity classifiers whenever those context records are missing or stale.

Which reliability decisions require human approval?▼

Approval is required for money movement, binding commitments, customer-impacting changes, production or safety changes, access changes, and actions above budget or risk limits. The skill executes only authorized low-risk reversible internal actions.