What problem does it solve?
This Skill helps SRE and platform engineers take ownership of an unfamiliar production service without guessing, rushing, or missing critical operational risk. It turns a handover into a structured safe-entry process that clarifies ownership, runtime behavior, dependencies, failure modes, and the next safest action.
Core Features & Use Cases
- Service understanding: Builds a clear overview of what the service does, who owns it, who depends on it, and what remains unknown.
- Runtime and infrastructure mapping: Traces traffic flow, deployment path, state, secrets, IaC, and blast radius across related systems.
- Risk and operations readiness: Surfaces critical workflows, incident history, observability gaps, runbook coverage, and safe first improvements or escalation points.
- Use case: A new on-call engineer can use this Skill to quickly understand a production service, identify the highest-risk paths, and decide whether to make a small safe change or document a defensible "do not touch yet" recommendation.
Quick Start
Use the service-reliability-onboarding skill to analyze this production service, map its dependencies and risks, and recommend the safest next step.