What problem does it solve?
This Skill addresses the complexity of building, optimizing, and maintaining robust, production-grade data systems and pipelines, ensuring scalability, reliability, and performance.
Core Features & Use Cases
- Data Pipeline Orchestration: Automate the execution, scheduling, and dependency management of complex ETL/ELT workflows using tools like Airflow, Prefect, or Dagster.
- Data Quality Validation: Implement rigorous checks for schema, constraints, and statistical anomalies to ensure data integrity.
- ETL Performance Optimization: Analyze and tune data processing jobs for maximum efficiency and minimal resource consumption.
- Use Case: A company needs to ingest, transform, and load terabytes of daily sales data into a data warehouse. This Skill can orchestrate the entire process, validate data quality at each stage, and optimize the transformation jobs for speed and cost-effectiveness.
Quick Start
Use the senior-data-engineer skill to orchestrate data pipelines starting from the 'data/' directory and outputting to 'results/'.