data-engineer

Designs scalable data pipelines and warehouses using Spark, dbt, Airflow, and cloud platforms.

1|1|Updated Feb 19, 2026
One-click install
npx skills add https://github.com/Dbillionaer/wholesaile --skill data-engineer-dbillionaer
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-engineer
Source: https://github.com/Dbillionaer/wholesaile/tree/main/skills/data-engineer
Command: npx skills add https://github.com/Dbillionaer/wholesaile --skill data-engineer-dbillionaer

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill addresses the complexity of building and maintaining robust, scalable data infrastructure, enabling efficient data processing and analytics.

Core Features & Use Cases

  • Data Pipeline Design: Architectures for batch and streaming data.
  • Modern Data Stack Implementation: Integration of tools like Spark, dbt, Airflow, and cloud platforms.
  • Use Case: Design and implement a real-time data pipeline to ingest, transform, and load streaming event data from Kafka into a Snowflake data warehouse for immediate business intelligence.

Quick Start

Design a scalable data pipeline to ingest streaming data from Kafka into Snowflake.

Frequently Asked Questions about data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I build a scalable data pipeline to ingest streaming data from Kafka into a data warehouse?

To build a scalable data pipeline, you architect a real-time streaming workflow that ingests, transforms, and loads event data from Kafka into platforms like Snowflake for immediate business intelligence.

What is the best way to design batch and stream processing architectures for large-scale operations?

Designing batch and stream processing architectures involves using Apache Spark and Airflow to orchestrate data ingestion, transformation, and governance across scalable data lakehouse environments.

Can I use dbt and Airflow together for data transformation and pipeline orchestration?

Yes, you can integrate dbt for data transformation and Airflow for pipeline orchestration within a modern data stack to solve complex challenges in large-scale data operations and warehousing.

Does this approach support cloud-native data platforms and data lakehouse architectures?

Yes, this approach supports cloud-native data platforms and data lakehouse architectures, implementing solutions for batch processing, stream processing, and modern cloud data warehousing.

How do I implement data governance and orchestration for complex data ingestion workflows?

Implementing data governance and orchestration involves using Airflow to schedule and manage complex data ingestion and transformation workflows across your scalable data infrastructure.