What problem does it solve? Founders and operators often lack a structured way to keep critical technology systems dependable: service objectives are undefined, incidents recur, changes break production, and recovery is never tested. This Skill turns reliability work into an evidence-backed decision and operating plan tied to business constraints. ## Core Features & Use Cases - Service tiering and SLO definition: Classify services by criticality, set service level objectives, and instrument reliability signals with confidence-scored evidence. - Incident, problem, and change control: Manage incidents and recurring problems, control change failure rate, and test recovery and continuity procedures. - Reliability economics review: Quantify risk-adjusted value, downside loss, and constraint headroom for reliability investments, simulated in the Business Digital Twin. - Use Case: A founder asks how to reduce outages without exceeding cash and risk limits. The agent tiers services, sets SLOs, compares at least three feasible options against the counterfactual, and recommends the highest confidence-weighted plan with monitoring and escalation rules. ## Quick Start Use technology reliability management to define SLOs and an incident and change control plan for our critical systems within our cash and risk limits.