data-engineer

Design and implement scalable data pipelines with Spark, dbt, and Airflow.

Updated Apr 9, 2026
One-click install
npx skills add https://github.com/raksok/netherica --skill data-engineer-raksok
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-engineer
Source: https://github.com/raksok/netherica/tree/main/.agents/skills/data-engineer
Command: npx skills add https://github.com/raksok/netherica --skill data-engineer-raksok

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Data engineering teams need to design, implement, and maintain robust data pipelines and modern analytics architectures that scale with data volume and variety.

Core Features & Use Cases

  • Modern data stack design and implementation across batch and streaming workloads
  • Data warehousing, lakehouse, and governance with quality controls
  • Cloud-native orchestration and deployment using Spark, dbt, and Airflow

Quick Start

Instantiate a scalable data pipeline using Spark, dbt, and Airflow to ingest, transform, and validate data in your data warehouse.

Frequently Asked Questions about data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I build scalable data pipelines using Spark, dbt, and Airflow?

Yes, modern data architecture supports both batch and streaming ETL workloads. You can design and implement scalable pipelines that process real-time analytics while applying data quality checks and governance controls across your cloud environment.

What is the best way to design a lakehouse architecture with data governance?

Yes, you can deploy cost-aware data pipelines on cloud platforms. The architecture applies monitoring to your ETL workflows, ensuring scalable data processing while optimizing operational costs across batch and streaming workloads.

Does dbt work with Airflow for data quality checks?

Yes, dbt works with Airflow to orchestrate data quality checks. You can integrate dbt transformations within Airflow pipelines to validate data and enforce governance controls during batch and streaming ETL processing.

How do I set up a modern data stack for real-time analytics?

You set up a modern data stack for real-time analytics by automating scalable data pipeline design. This integrates Spark, dbt, and Airflow to ingest, transform, and validate streaming data in your cloud data warehouse.