data-engineering

Automates design, implementation, and validation of ETL pipelines and data warehouses.

2.5k|877|Updated Feb 4, 2026
One-click install
npx skills add https://github.com/rohitg00/awesome-claude-code-toolkit --skill data-engineering-rohitg00
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-engineering
Source: https://github.com/rohitg00/awesome-claude-code-toolkit/tree/main/skills/data-engineering
Command: npx skills add https://github.com/rohitg00/awesome-claude-code-toolkit --skill data-engineering-rohitg00

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Data engineers often struggle to design, implement, and maintain scalable ETL pipelines and data warehouses with reliable quality checks.

Core Features & Use Cases

  • Pattern templates for ETL pipelines, data warehousing schemas, Spark processing, and data quality validation to accelerate implementation.
  • Use Case: rapidly prototype end-to-end data workflows from ingestion to analytics with built-in validation and monitoring.

Quick Start

Create a starter data-engineering workflow skeleton using the included templates for an end-to-end ETL pipeline.

Frequently Asked Questions about data-engineering

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design and validate scalable ETL pipelines?

To design and validate scalable ETL pipelines, use pattern templates for extraction, transformation, loading, and quality validation steps. These templates provide built-in monitoring to accelerate the implementation of robust data workflows from ingestion to analytics.

What is the best way to build a star-schema data warehouse?

The best way to build a star-schema data warehouse is using pre-defined pattern templates that satisfy data warehousing schema requirements. This approach automates schema design and integrates seamlessly with ETL pipelines for reliable analytics.

Can I use Spark for data processing and quality validation in my ETL workflow?

Yes, you can use Spark for data processing and quality validation in your ETL workflow. The patterns include Spark-based analytics examples and validation logic to ensure data quality across both cloud and on-prem environments.

How do I add data quality checks to an existing ETL pipeline?

To add data quality checks to an existing ETL pipeline, apply the included data validation pattern templates. These templates provide SQL examples and validation logic to implement reliable monitoring and quality control across your data workflows.

Does this ETL pipeline approach work for both cloud and on-prem environments?

Yes, this ETL pipeline approach works for both cloud and on-prem environments. The design patterns and validation templates are built to support scalable data workflows across diverse infrastructure setups without environmental limitations.