data-system-ops-lead

Automate data system operations and reliability governance for data platforms.

7|1|Updated May 19, 2026
One-click install
npx skills add https://github.com/daemon-blockint-tech/Agentic-Enteprises-Skill --skill data-system-ops-lead
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-system-ops-lead
Source: https://github.com/daemon-blockint-tech/Agentic-Enteprises-Skill/tree/main/data-system-ops-lead
Command: npx skills add https://github.com/daemon-blockint-tech/Agentic-Enteprises-Skill --skill data-system-ops-lead

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Run data system operations and reliability engineering. Cover pipeline monitoring, incident response, SLA management, capacity planning, on-call runbooks, data quality alerting, and operational excellence. Triggers on "data pipeline monitoring", "incident response", "SLA management", "capacity planning", "on-call runbook", "data quality alerting", "operational excellence", "system reliability", "pipeline health check", or "data ops".

Core Features & Use Cases

  • Pipeline monitoring with alerting thresholds and dashboard design
  • Incident response: severity classification, escalation paths, post-incident reviews
  • SLA management with performance tracking and breach prevention
  • Capacity planning: resource forecasting, scaling triggers, cost optimization
  • On-call runbooks with step-by-step troubleshooting procedures
  • Data quality alerting with anomaly detection and validation rules
  • Operational excellence and governance across data platforms
  • Use Case: When a data platform experiences lag, this skill guides the ops workflow to restore availability and meet SLAs.

Quick Start

Run a daily health check and open the on-call runbook to start the workflow.

Frequently Asked Questions about data-system-ops-lead

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I set up incident response and escalation paths for data pipeline monitoring?

Data pipeline monitoring incident response requires structured runbooks with severity classification and defined escalation paths. This skill guides the workflow to restore availability, conduct health checks, and perform post-incident reviews to maintain SLA compliance.

What is the best way to manage SLA breaches and track data platform reliability?

SLA management involves performance tracking and breach prevention across data platforms. This skill automates reliability governance by aligning operational procedures with capacity management, data quality alerting, and standard change control workflows.

How do I create on-call runbooks for data quality alerting and anomaly detection?

On-call runbooks for data quality alerting provide step-by-step troubleshooting procedures triggered by anomaly detection and validation rules. This skill supports end-to-end workflows including health checks, structured runbooks, and operational excellence governance.

Can I use this for capacity planning and resource forecasting on data platforms?

Capacity planning for data platforms includes resource forecasting, scaling triggers, and cost optimization. This skill aligns capacity management with vendor management and operational procedures to support data system operations at scale.

Why do I need operational excellence governance for data ops and what does it cover?

Operational excellence governance covers pipeline monitoring, incident response, SLA management, and capacity planning across data platforms. This skill establishes standard operational procedures for change control, vendor management, and continuous reliability engineering.