data-engineer

Automate design and deployment of scalable ETL/ELT and streaming data pipelines.

1|Updated Sep 11, 2025
One-click install
npx skills add https://github.com/Dhumitech/DHUMI-AI-RESOURCE --skill data-engineer-dhumitech
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-engineer
Source: https://github.com/Dhumitech/DHUMI-AI-RESOURCE/tree/main/AI-Engineer-planner-Skills/02-data/data-engineer
Command: npx skills add https://github.com/Dhumitech/DHUMI-AI-RESOURCE --skill data-engineer-dhumitech

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Design and build data pipelines across ETL, data validation, schema evolution, and streaming architectures to ensure scalable, reliable analytics.

Core Features & Use Cases

  • End-to-end data pipeline design and orchestration
  • Batch and streaming processing with modern data stack (dbt, Airflow, Spark)
  • Data quality, governance, and monitoring for production pipelines

Quick Start

Define sources, schemas, and SLAs, then deploy a minimal end-to-end data pipeline to validate the workflow.

Frequently Asked Questions about data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design scalable data pipelines for batch and streaming processing?

Scalable data pipelines are designed by automating the deployment of robust ETL/ELT and streaming architectures. This approach uses a modern data stack including dbt, Airflow, and Spark to handle ingestion, validation, and monitoring for production analytics.

What is the best way to orchestrate ETL workflows with governance and data quality checks?

Orchestrating ETL workflows with governance requires automating pipeline design and applying quality checks across the data platform. Using tools like Airflow and dbt ensures reliable analytics, schema evolution, and cost-aware monitoring throughout the data lifecycle.

Do I need a modern data stack like dbt and Spark to build reliable data pipelines?

A modern data stack including dbt, Airflow, and Spark is required to build reliable data pipelines. This Skill assumes your organization has tooling for ingestion, validation, and monitoring to deploy robust batch and streaming architectures effectively.

Can I use Airflow for streaming data architectures alongside batch ETL?

Airflow orchestrates both batch and streaming data architectures within a unified data platform. It coordinates the automated deployment of scalable pipelines, ensuring data quality and governance checks are applied consistently across all processing workflows.

How do I start building a minimal end-to-end data pipeline to validate my workflow?

To start building a data pipeline, define your data sources, schemas, and SLAs first. You can then deploy a minimal end-to-end pipeline to validate the workflow before scaling up to full production environments with governance and monitoring.

Why does my data pipeline require schema evolution and data validation monitoring?

Data pipeline schema evolution and validation monitoring are required to ensure reliable analytics and maintain data quality in production. Applying governance and automated checks prevents downstream errors and supports scalable, cost-aware orchestration across the platform.