ai-ops

Automate runbook execution, incident response, and operational health monitoring.

54|3|Updated Feb 4, 2026
One-click install
npx skills add https://github.com/arcasilesgroup/ai-engineering --skill ai-ops-arcasilesgroup
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ai-ops
Source: https://github.com/arcasilesgroup/ai-engineering/tree/main/.claude/skills/ai-ops
Command: npx skills add https://github.com/arcasilesgroup/ai-engineering --skill ai-ops-arcasilesgroup

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill streamlines operational tasks, automates incident response, and provides real-time monitoring of system health, reducing manual intervention and improving reliability.

Core Features & Use Cases

  • Runbook Execution: Automates the execution of predefined operational runbooks for routine tasks and complex procedures.
  • Incident Management: Diagnoses failures, classifies severity, and initiates recovery steps for incidents.
  • Operational Status Monitoring: Aggregates signals to provide a clear dashboard of system health.
  • Use Case: When a CI pipeline fails, this Skill can automatically diagnose the issue, identify the relevant runbook, and attempt to fix it, or escalate if necessary.

Quick Start

Use the ai-ops skill to check the operational status of the system.

Frequently Asked Questions about ai-ops

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate incident response and runbook execution for system failures?

This Skill streamlines operational tasks by automating runbook execution, diagnosing failures, classifying incident severity, and proposing or executing recovery steps to reduce manual intervention.

What is the best way to monitor operational health for AI-governed systems?

To monitor operational health, this Skill aggregates signals from runbook execution success rates, decision store status, CI pipeline health, and issue backlog to provide a clear system status dashboard.

How does automated runbook execution handle CI pipeline failures?

When a CI pipeline fails, the Skill automatically diagnoses the issue, identifies the relevant runbook, attempts to fix it, and escalates the incident if automated recovery is not possible.

Can I use this for diagnosing failures and classifying incident severity?

Yes, this Skill diagnoses system failures and classifies incident severity to propose and execute recovery steps, managing reliability for AI-governed systems throughout the incident lifecycle.

Do I need predefined operational runbooks to automate incident management?

Yes, predefined operational runbooks are required, as the Skill automates incident management and routine tasks by executing the specific runbooks established for your operational procedures.