What problem does it solve?
This Skill reduces incidents, downtime, and runaway cloud costs by providing a systematic, automated approach to reliability, observability, backup, and compliance for production infrastructure.
Core Features & Use Cases
- Monitoring & Alerting: Prometheus-based metric collection and alert rules to detect CPU, memory, disk, and service outages before they impact users.
- Infrastructure as Code: Terraform patterns for VPCs, subnets, autoscaling groups, and managed databases to enable repeatable, auditable deployments.
- Backup & Recovery: Automated backup scripts with encryption, S3 uploads, verification, and retention policies to ensure recoverability.
- Security & Compliance: Guidance for access control, patching, vulnerability management, and audit trails to meet regulatory requirements.
- Use Case: Use this Skill to plan a migration to multi-AZ autoscaling with Prometheus alerts, implement automated backups with encrypted S3 storage, and produce a post-change rollback plan and runbook.
Quick Start
Use the Infrastructure Maintainer to assess your current cloud environment, generate prioritized optimizations and Terraform changes, and produce a tested backup and incident response playbook.