data-engineer

Design batch and streaming data pipelines with Spark, dbt, and Airflow.

Updated Dec 10, 2024
One-click install
npx skills add https://github.com/melikhanmutlu/web_ar --skill data-engineer-melikhanmutlu
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-engineer
Source: https://github.com/melikhanmutlu/web_ar/tree/main/skills-extra/data-engineer
Command: npx skills add https://github.com/melikhanmutlu/web_ar --skill data-engineer-melikhanmutlu

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Data engineering teams struggle to design, implement, and maintain robust data pipelines that scale with growing data volumes, integrate diverse sources, enforce data quality, and provide reliable analytics infrastructure.

Core Features & Use Cases

  • End-to-end pipeline design for batch and streaming workloads using Spark, dbt, and Airflow.
  • Modern data warehousing and lakehouse architectures with governance and quality controls.
  • Data quality, lineage, and observability across the analytics stack; suitable for ETL/ELT and data products.

Quick Start

Outline how to design, implement, and operate scalable data pipelines and modern analytics platforms.

Frequently Asked Questions about data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design scalable data pipelines for batch and streaming workloads?

Design scalable data pipelines by specifying data sources, orchestration, ingestion, and validation. Apply frameworks like Spark, dbt, and Airflow to build robust infrastructure for both batch and streaming workloads.

What is the best way to enforce data quality and observability across an analytics stack?

Enforce data quality and observability by integrating validation and monitoring controls into your pipelines. Ensure lineage and quality checks are applied across the analytics stack for reliable data products.

Can I use this approach to build a modern lakehouse architecture with governance?

Yes, you can build modern data warehousing and lakehouse architectures. The approach includes specifying governance, quality controls, and cost guardrails across cloud platforms.

How do I orchestrate ETL and ELT workflows using Airflow and dbt?

Orchestrate ETL and ELT workflows by integrating dbt for transformations and Airflow for scheduling. This combination supports end-to-end pipeline design from data ingestion to analytics product delivery.

When do I need Great Expectations for data pipeline validation?

You need Great Expectations for data pipeline validation when enforcing strict data quality and governance. It provides the necessary controls for monitoring, lineage tracking, and observability across your infrastructure.