senior-data-engineer

Design and implement scalable data pipelines across Python, SQL, Spark, Airflow, dbt, and Kafka.

1|Updated Sep 16, 2025
One-click install
npx skills add https://github.com/zhizhunbao/ottawa-genai-research-assistant --skill senior-data-engineer-zhizhunbao
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: senior-data-engineer
Source: https://github.com/zhizhunbao/ottawa-genai-research-assistant/tree/main/.agent/skills/dev-senior_data_engineer
Command: npx skills add https://github.com/zhizhunbao/ottawa-genai-research-assistant --skill senior-data-engineer-zhizhunbao

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

Data engineering challenges: designing scalable pipelines, ETL/ELT orchestration, and governance across modern stacks.

Core Features & Use Cases

  • Production-grade pipeline design, orchestration, and data quality across Python, SQL, Spark, Airflow, dbt, and Kafka.
  • Data modeling, pipeline orchestration, and DataOps governance for end-to-end analytics platforms.
  • Use cases include building reliable data lakes, real-time streaming pipelines, and scalable data platforms for enterprises.

Quick Start

Describe an end-to-end data engineering workflow from ingestion to deployment, including orchestration and data quality checks.

Frequently Asked Questions about senior-data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design scalable data pipelines for enterprise data platforms?

Design scalable data pipelines by applying architecture design, ETL/ELT orchestration, and data modeling across Python, SQL, Spark, Airflow, dbt, and Kafka to achieve production-grade reliability and scalable data operations.

What's the best way to orchestrate ETL workflows with Airflow and dbt?

Orchestrate ETL workflows by integrating Airflow for pipeline scheduling and dbt for data modeling and transformations, ensuring end-to-end execution, observability, and DataOps governance across analytics platforms.

How do I implement data quality and governance checks in Spark data pipelines?

Implement data quality and governance in Spark pipelines through validation rules, observability monitoring, and DataOps governance frameworks to maintain production-grade reliability and accurate analytics outputs.

Can I build real-time streaming pipelines with Kafka and Python for data lakes?

Build real-time streaming pipelines using Kafka and Python to ingest data into scalable data lakes, enabling enterprise data platforms with production-grade reliability and continuous data operations.

When do I need to use dbt for data modeling versus raw SQL transformations?

Use dbt for data modeling when orchestrating scalable ELT workflows with built-in governance and data quality validation, whereas raw SQL suits isolated transformations lacking pipeline observability and orchestration requirements.

Why does my data pipeline orchestration fail to scale for enterprise data platforms?

Data pipeline orchestration fails to scale without proper architecture design, distributed processing via Spark, and robust DataOps governance, leading to bottlenecks in ETL/ELT workflows and degraded observability.