senior-data-engineer

Design and operate scalable data pipelines with Python, SQL, Spark, Airflow, dbt, and Kafka.

Updated Mar 21, 2026
One-click install
npx skills add https://github.com/AgLyx3/My-Note-App-Not-Just-a-Note-App --skill senior-data-engineer-aglyx3
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: senior-data-engineer
Source: https://github.com/AgLyx3/My-Note-App-Not-Just-a-Note-App/tree/main/.cursor/skills/engineering-team/senior-data-engineer
Command: npx skills add https://github.com/AgLyx3/My-Note-App-Not-Just-a-Note-App --skill senior-data-engineer-aglyx3

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires yaml, and includes scripts (resource) and references (resource) components.

What problem does it solve?

Data engineers often struggle to design, build, and operate scalable data pipelines and infrastructure that can handle batch and streaming workloads, data quality, governance, and operability across teams.

Core Features & Use Cases

  • End-to-end data pipeline design and implementation across Python, SQL, Spark, Airflow, dbt, and Kafka stacks.
  • Advanced data modeling, pipeline orchestration, data quality enforcement, and DataOps practices for reliable analytics and governance.
  • Use case: architect a lakehouse pipeline with streaming and batch components, monitor quality, and troubleshoot performance issues in production.

Quick Start

Describe a ready-to-run data engineering blueprint for a lakehouse pipeline, including tooling, schema, and governance considerations.

Frequently Asked Questions about senior-data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design scalable data pipelines for batch and streaming workloads?

Scalable data pipelines for batch and streaming workloads are designed by integrating Spark, Kafka, and Airflow to orchestrate ETL processes, ensuring reliable data infrastructure and analytics across Python and SQL stacks.

What is the best way to architect a lakehouse pipeline with streaming and batch components?

Architecting a lakehouse pipeline with streaming and batch components involves using Spark for processing, Kafka for streaming ingestion, and dbt for data modeling, while enforcing data quality and governance throughout the workflow.

How do I implement DataOps practices for monitoring data quality in ETL systems?

Implementing DataOps for monitoring data quality in ETL systems requires orchestrating pipelines with Airflow, applying data modeling via dbt, and enforcing governance checks to maintain reliable analytics and operational health.

Does this data engineering approach support troubleshooting performance issues in production?

Yes, this data engineering approach supports troubleshooting production performance issues by leveraging pipeline orchestration, DataOps practices, and infrastructure monitoring across Python, SQL, Spark, and Kafka environments.

When should I use Airflow and dbt for pipeline orchestration and data modeling?

Use Airflow and dbt for pipeline orchestration and data modeling when building scalable ETL systems that require strict data quality enforcement, reliable workflow automation, and structured governance across modern data stacks.