astra-sre

Orchestrate health scans, incident triage, guided repair, and learning loops for multi-node infrastructure.

1|Updated Jun 18, 2026
One-click install
npx skills add https://github.com/alrcatraz/astra-aiagent-infra --skill astra-sre
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: astra-sre
Source: https://github.com/alrcatraz/astra-aiagent-infra/tree/main/skills/devops/astra-sre
Command: npx skills add https://github.com/alrcatraz/astra-aiagent-infra --skill astra-sre

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires astra-hub, astra-sre-fix-e2ee, astra-sre-restart-service, astra-sre-fix-gfw, astra-sre-fix-mcp, astra-sre-fix-vps-recovery, server-restart-recovery, server-health-audit, infrastructure-device-inventory, service-inventory, crash-marker-pattern, full-e2ee-recovery-after-server-rebuild, pre-upgrade-server-backup, and includes scripts (resource) and references (resource) components.

What problem does it solve?

The astra-sre skill provides a unified SRE (Site Reliability Engineering) coordination layer for multi-node infrastructure, addressing the challenges of health scanning, incident triage, guided repair, and learning across all managed devices.

Core Features & Use Cases

  • Health Scanning: Conducts comprehensive health scans across all devices and services.
  • Incident Triage: Assesses the severity and impact of incidents, routing known faults to appropriate sub-skills.
  • Guided Repair: Offers a step-by-step repair plan based on diagnosis, with verification probes and rollback mechanisms.
  • Learning Loop: Continuously learns from incidents and suggests creating new sub-skills for recurring issues.
  • Use Case: When an incident occurs, astra-sre can automatically diagnose the problem, propose a repair plan, and guide the user through the process to ensure a reliable infrastructure.

Quick Start

To begin, activate astra-sre to initiate a health scan of your infrastructure.

Frequently Asked Questions about astra-sre

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I coordinate incident triage and health scanning for multi-node infrastructure?

Infrastructure health scanning and incident triage are coordinated by assessing incident severity, routing known faults to sub-skills, and conducting comprehensive health checks across all managed devices and services.

What is the best way to automate guided repair and fault diagnostics for reliability engineering?

Guided repair for reliability engineering is automated by offering a step-by-step repair plan based on fault diagnostics, complete with verification probes and rollback mechanisms to ensure stable multi-node infrastructure.

How does a site reliability engineering learning loop handle recurring infrastructure incidents?

A site reliability engineering learning loop handles recurring incidents by continuously learning from fault diagnostics and suggesting the creation of new sub-skills for known issues to prevent future infrastructure failures.

Do I need pre-upgrade server backups and device inventory data to use SRE incident management?

Yes, SRE incident management requires pre-upgrade server backups, device inventory data, and service inventory integration to accurately assess fault severity and route incidents to appropriate repair sub-skills.

Can I use automated SRE fixes for full e2ee recovery after a server rebuild?

Yes, automated SRE fixes can handle full e2ee recovery after a server rebuild by orchestrating guided repair plans and utilizing specialized sub-skills like e2ee recovery and server restart mechanisms.

What are the limitations of using a unified SRE coordination layer for multi-node infrastructure management?

The limitations of unified SRE coordination include its dependency on multiple sub-skills for specific fixes like gfw or mcp issues, meaning it routes faults rather than independently resolving all multi-node infrastructure anomalies.