data-engineer

Automate scalable data pipeline design and management with Apache Spark, dbt, and Airflow.

Updated Apr 17, 2026
One-click install
npx skills add https://github.com/CompSci-Squad/tcc_ai --skill data-engineer-compsci-squad
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-engineer
Source: https://github.com/CompSci-Squad/tcc_ai/tree/main/.github/skills/data-engineer
Command: npx skills add https://github.com/CompSci-Squad/tcc_ai --skill data-engineer-compsci-squad

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires Apache Spark, dbt, Airflow, Fivetran/Airbyte, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill addresses the complexity of building and managing scalable data pipelines and architectures, offering expertise in modern data technologies and best practices.

Core Features & Use Cases

  • Data Pipeline Design: Guides in designing scalable and efficient data pipelines.
  • Data Architecture Implementation: Assists in setting up modern data architectures like data lakehouses and cloud data warehouses.
  • Use Case: Build a robust and scalable data pipeline from scratch, integrating multiple data sources, implementing transformations, and validating the data flow.

Quick Start

Use the data-engineer skill to architect a scalable data pipeline that connects various data sources, including databases, files, and APIs.

Frequently Asked Questions about data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I build a scalable data pipeline from scratch?

To build a scalable data pipeline, you must design data architecture that integrates multiple sources, implements transformations, and validates data flow. This Skill automates that design process using Apache Spark, dbt, and Airflow to ensure robust enterprise data processing.

What is the best way to orchestrate ETL/ELT workflows in a cloud data platform?

Orchestrating ETL/ELT workflows requires managing dependencies and scheduling data transformations. This Skill guides you in using Airflow for orchestration and dbt for transformations within cloud data platforms to automate scalable data integration.

Can I use Apache Spark and dbt together for data architecture implementation?

Yes, you can use Apache Spark and dbt together for data architecture implementation. Spark handles large-scale data processing, while dbt manages the SQL transformations, allowing you to build scalable data lakehouses and cloud data warehouses efficiently.

Do I need to know cloud data platforms to automate data pipeline design?

Yes, automating data pipeline design in this context requires pre-existing knowledge of cloud data platforms, data integration tools, and data orchestration. The Skill applies to enterprise environments where you must manage robust data processing and storage solutions.

How does Fivetran or Airbyte fit into modern data architecture?

Fivetran and Airbyte fit into modern data architecture by automating data extraction and loading from various sources into your data platform. This Skill incorporates these tools to connect databases, files, and APIs, streamlining the initial stages of your data pipeline.

When should I not use a data lakehouse architecture for enterprise data processing?

You should not use a data lakehouse architecture for enterprise data processing if your environment lacks the scalable infrastructure needed for Apache Spark or cloud data platforms. This Skill targets complex, robust data pipelines requiring significant orchestration and integration capabilities.