What problem does it solve?
IT operations teams struggle with alert fatigue, inconsistent incident response, manual repetitive tasks, and untested disaster recovery plans. This Skill provides structured frameworks, runbooks, and decision matrices to run reliable infrastructure operations.
Core Features & Use Cases
- Monitoring & Observability: Guidance on SLI/SLO/SLA definitions, the Four Golden Signals, alert threshold tuning, and dashboard design.
- Incident Management: Severity classification (P1-P4), incident response roles, communication templates, root cause analysis (5 Whys, fishbone), and blameless post-mortem templates.
- Automation & Toil Reduction: ROI calculation for automation, Bash/Python scripting patterns, Ansible playbooks, and GitOps CI/CD pipelines for infrastructure.
- Backup & Disaster Recovery: 3-2-1 backup rule, RPO/RTO planning, DR site types, recovery testing checklists, and database backup scripts.
- Use Case: When a P1 outage hits production, use this Skill to classify severity, assign incident roles, follow the response workflow, communicate with stakeholders, and produce a blameless post-mortem with tracked action items.
Quick Start
Ask the assistant to help you design an incident response plan with severity levels and on-call escalation procedures for your production services.