agency-infrastructure-maintainer

Maintain cloud infrastructure reliability, performance, and security with Terraform and Prometheus.

Updated Feb 11, 2026
One-click install
npx skills add https://github.com/augustoheiss/LogicDefense --skill agency-infrastructure-maintainer-augustoheiss
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: agency-infrastructure-maintainer
Source: https://github.com/augustoheiss/LogicDefense/tree/main/.gemini/skills/agency-infrastructure-maintainer
Command: npx skills add https://github.com/augustoheiss/LogicDefense --skill agency-infrastructure-maintainer-augustoheiss

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill reduces incidents, downtime, and runaway cloud costs by providing a systematic, automated approach to reliability, observability, backup, and compliance for production infrastructure.

Core Features & Use Cases

  • Monitoring & Alerting: Prometheus-based metric collection and alert rules to detect CPU, memory, disk, and service outages before they impact users.
  • Infrastructure as Code: Terraform patterns for VPCs, subnets, autoscaling groups, and managed databases to enable repeatable, auditable deployments.
  • Backup & Recovery: Automated backup scripts with encryption, S3 uploads, verification, and retention policies to ensure recoverability.
  • Security & Compliance: Guidance for access control, patching, vulnerability management, and audit trails to meet regulatory requirements.
  • Use Case: Use this Skill to plan a migration to multi-AZ autoscaling with Prometheus alerts, implement automated backups with encrypted S3 storage, and produce a post-change rollback plan and runbook.

Quick Start

Use the Infrastructure Maintainer to assess your current cloud environment, generate prioritized optimizations and Terraform changes, and produce a tested backup and incident response playbook.

Frequently Asked Questions about agency-infrastructure-maintainer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I set up Prometheus monitoring and alerting for cloud infrastructure?

To set up Prometheus monitoring for cloud infrastructure, you need metric collection pipelines and alert rules configured to detect CPU, memory, disk, and service outages before they impact users. This Skill provides the systematic patterns to establish these observability configurations.

What is the best way to automate cloud infrastructure backups with S3 storage?

Automating cloud infrastructure backups with S3 requires scripts that handle encryption, S3 uploads, verification, and retention policies. This approach ensures data recoverability and compliance for production environments by automating the entire backup and recovery process.

How do I use Terraform to manage multi-AZ autoscaling deployments?

Using Terraform for multi-AZ autoscaling deployments involves applying patterns for VPCs, subnets, autoscaling groups, and managed databases. These infrastructure as code patterns enable repeatable, auditable deployments across multi-region cloud environments.

Can I generate an incident response runbook for cloud-native applications?

Yes, you can generate an incident response runbook for cloud-native applications. This Skill produces operational runbooks covering incident response, monitoring, backup, and cost optimization to maintain system reliability and security during production outages.

How do I implement compliance-driven access controls for production infrastructure?

Implementing compliance-driven access controls for production infrastructure requires guidance for access control, patching, vulnerability management, and audit trails. This Skill provides the systematic framework to meet regulatory requirements and maintain security.

Does this infrastructure maintenance approach work for multi-region deployments?

Yes, this infrastructure maintenance approach works for multi-region deployments. It is explicitly applicable to multi-region deployments and cloud-native applications, providing automated approaches to reliability, observability, backup, and compliance across distributed environments.