data-engineer

Design scalable data pipelines and warehouses with Apache Spark, dbt, and Airflow.

Updated Mar 11, 2026
One-click install
npx skills add https://github.com/Industrial/rust-symphony --skill data-engineer-industrial
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-engineer
Source: https://github.com/Industrial/rust-symphony/tree/main/.cursor/skills/data-engineer
Command: npx skills add https://github.com/Industrial/rust-symphony --skill data-engineer-industrial

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill addresses the complexity of designing, building, and maintaining robust, scalable data pipelines and modern data architectures.

Core Features & Use Cases

  • Pipeline Design & Implementation: Architect batch and streaming data pipelines, data warehouses, and lakehouses.
  • Technology Integration: Leverages a wide array of tools including Apache Spark, dbt, Airflow, and cloud-native platforms.
  • Use Case: Design a real-time streaming pipeline that processes 1M events per second from Kafka to BigQuery.

Quick Start

Use the data-engineer skill to design a modern data stack with dbt, Snowflake, and Fivetran for dimensional modeling.

Frequently Asked Questions about data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I build scalable data pipelines for both batch and streaming processing?

You build scalable data pipelines using Apache Spark, Airflow, and cloud-native data platforms to handle both batch and streaming processing. This architecture solves complex pipeline design problems by orchestrating robust data workflows across modern data stacks.

What is the best way to architect a real-time streaming pipeline with Kafka and BigQuery?

Architecting a real-time streaming pipeline from Kafka to BigQuery involves leveraging cloud-native platforms and Spark. This design processes high-throughput events, scaling to handle 1M events per second for real-time analytics infrastructure.

How do I design a modern data warehouse using dbt and Snowflake?

Design a modern data warehouse using dbt and Snowflake by implementing dimensional modeling within a modern data stack. This integrates cloud-native platforms to structure analytics infrastructure, enabling scalable data warehousing and lakehouse architectures.

When do I need a lakehouse architecture instead of a standard data warehouse?

You need a lakehouse architecture when unifying batch and streaming processing within scalable data pipelines. This approach combines data warehousing and lakehouse designs using Apache Spark and cloud services to solve complex analytics infrastructure requirements.

Can I use Apache Airflow to orchestrate dbt transformations in a data pipeline?

Yes, Apache Airflow orchestrates dbt transformations within scalable data pipelines to automate analytics infrastructure. This integration leverages modern data stack implementation, coordinating batch processing and data warehousing workflows effectively.