What problem does it solve?
This Skill eliminates the guesswork and operational risk of managing cloud infrastructure, CI/CD pipelines, runtime services, and production systems for engineering teams, reducing rollout failures, incident response time, and reliability gaps without requiring deep specialized expertise.
Core Features & Use Cases
- Infrastructure & IaC Best Practices: Provides guidance for Terraform, GCP, Ansible, and container deployments, including state management, drift resolution, safe change processes, and rollback planning.
- Runtime & Incident Response: Helps troubleshoot production issues like 502 errors, runtime saturation, and deployment failures with actionable debugging steps, blast radius assessment, and safe remediation paths.
- Use Case Example: If you are preparing to replace a Cloud SQL instance via Terraform and need to validate IAM changes, firewall rules, and rollback steps before applying the plan, this skill walks you through the full safety review process.
Quick Start
Use the senior-devops-engineer skill to review your Terraform plan for a Cloud SQL replacement and IAM firewall changes, then get a rollout safety assessment with explicit rollback steps.