data-engineering-pro

Orchestrate ETL/ELT pipelines and data warehouse architectures with Kafka, Spark, and Flink.

Updated Jun 27, 2026
One-click install
npx skills add https://github.com/truongnat/aix --skill data-engineering-pro-truongnat
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-engineering-pro
Source: https://github.com/truongnat/aix/tree/main/content/skills/data-engineering-pro
Command: npx skills add https://github.com/truongnat/aix --skill data-engineering-pro-truongnat

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill addresses the challenge of building and managing complex data pipelines and storage architectures, providing expertise in ETL/ELT, distributed processing, and data warehouse design.

Core Features & Use Cases

  • ETL/ELT Pipelines: Construct robust ETL or ELT pipelines for analytical databases.
  • Data Warehouse Architecture: Design data warehouse or data lake architectures.
  • Event Streaming: Implement real-time event streaming architectures with Kafka.
  • Data Transformation: Optimize slow data transformation jobs using SQL or PySpark.

Quick Start

Analyze and optimize your data pipeline with the data-engineering-pro skill.

Frequently Asked Questions about data-engineering-pro

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design a scalable data warehouse architecture for analytical databases?

Designing a scalable data warehouse architecture requires expert orchestration of storage structures and ETL pipelines. This Skill provides architectural blueprints for data lakes and warehouses, ensuring robust analytical database construction.

What is the best way to optimize slow data transformation jobs in PySpark?

Optimizing slow data transformation jobs involves refining PySpark and SQL execution plans. This Skill analyzes pipeline bottlenecks and provides expert-level optimizations to accelerate distributed processing tasks.

How does Kafka event streaming integrate into a real-time data pipeline?

Kafka event streaming integrates into real-time pipelines by handling continuous data ingestion and orchestration. This Skill implements robust streaming architectures to process event flows efficiently.

Can I build ETL pipelines for distributed processing environments like Spark and Flink?

You can build ETL pipelines for distributed processing using Spark and Flink. This Skill orchestrates scalable data ingestion and transformation tasks tailored for distributed computing frameworks.

When do I need an ELT pipeline instead of an ETL pipeline for my data warehouse?

You need an ELT pipeline when transforming data directly within the target data warehouse or data lake rather than a separate staging area. This Skill structures both ETL and ELT workflows based on your storage architecture.

What are the limitations of building data lakes without structured orchestration?

Building data lakes without structured orchestration leads to unmanaged ingestion and inefficient storage architectures. This Skill mitigates such constraints by providing expert-level pipeline orchestration and data transformation handling.