data-engineer

Design and implement scalable data pipelines with Apache Spark, dbt, and Airflow.

5|2|Updated Jan 24, 2026
One-click install
npx skills add https://github.com/s1366560/agi-demos --skill data-engineer-s1366560
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-engineer
Source: https://github.com/s1366560/agi-demos/tree/main/.memstack/skills/data-engineer
Command: npx skills add https://github.com/s1366560/agi-demos --skill data-engineer-s1366560

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill addresses the complexity of designing, building, and maintaining robust, scalable data pipelines and modern data architectures for enterprise needs.

Core Features & Use Cases

  • Data Pipeline Design: Architect batch and streaming data pipelines.
  • Modern Data Stack Implementation: Integrate tools like Spark, dbt, Airflow, and cloud platforms.
  • Use Case: Implement a real-time streaming pipeline that processes 1 million events per second from Kafka to BigQuery, ensuring data quality and reliability.

Quick Start

Design a scalable data pipeline for processing real-time streaming data.

Frequently Asked Questions about data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design a scalable data pipeline for processing real-time streaming data?

To design a scalable data pipeline for real-time streaming data, architect batch and streaming workflows using Apache Spark and cloud-native platforms to process millions of events per second. This ensures data quality, governance, and performance optimization for high-throughput workloads.

What is the modern data stack and how do I implement it for enterprise data warehousing?

The modern data stack integrates tools like Spark, dbt, Airflow, and cloud platforms to build scalable data warehouses and robust data pipelines. Implementing it solves the complexity of maintaining enterprise data architectures while ensuring data lineage and governance.

Can I use Apache Airflow with dbt for batch and streaming data pipelines?

Yes, you can use Apache Airflow with dbt to orchestrate batch and streaming data pipelines. Integrating these modern data stack components allows you to schedule workflows, manage data transformations, and maintain data quality across both batch and streaming workloads.

How do I ensure data quality and lineage when processing 1 million streaming events per second from Kafka to BigQuery?

To ensure data quality and lineage when processing 1 million streaming events per second from Kafka to BigQuery, implement a real-time streaming architecture focused on data governance and performance optimization. This approach maintains reliability for high-volume data pipelines.

What is the best way to architect a real-time streaming pipeline for enterprise data needs?

The best way to architect a real-time streaming pipeline for enterprise data needs is to leverage cloud-native platforms and technologies like Apache Spark. This approach designs scalable data pipelines that guarantee data quality, reliability, and performance optimization for streaming workloads.