senior-data-engineer

Build scalable data pipelines with Python, SQL, Spark, Airflow, dbt, and Kafka.

27|10|Updated Dec 27, 2025
One-click install
npx skills add https://github.com/nilecui/SkillsBase --skill senior-data-engineer-nilecui
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: senior-data-engineer
Source: https://github.com/nilecui/SkillsBase/tree/main/.cursor/skills/senior-data-engineer
Command: npx skills add https://github.com/nilecui/SkillsBase --skill senior-data-engineer-nilecui

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill addresses the complex challenges of building and maintaining robust, scalable data pipelines, ETL/ELT systems, and data infrastructure for production-grade AI/ML/Data systems.

Core Features & Use Cases

  • Data Pipeline Orchestration: Automate and manage complex data workflows.
  • Data Quality Validation: Ensure the integrity and accuracy of data throughout the pipeline.
  • ETL/ELT Optimization: Enhance the performance and efficiency of data transformation processes.
  • Use Case: A company needs to ingest data from multiple sources, transform it, and load it into a data warehouse for business intelligence. This Skill can orchestrate the entire process, validate data quality at each step, and optimize the performance of the transformations.

Quick Start

Use the senior-data-engineer skill to orchestrate data pipelines using the provided scripts.

Frequently Asked Questions about senior-data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I build scalable data pipelines for production-grade AI/ML systems?

Build scalable data pipelines by orchestrating ETL/ELT workflows using Python, SQL, and Spark. This Skill automates data ingestion, transformation, and loading into data warehouses, optimizing performance and ensuring data quality for production-grade AI/ML systems.

What is DataOps and how does it apply to data pipeline orchestration?

DataOps applies to data pipeline orchestration by integrating data quality validation, governance, and automated workflow management. It ensures data integrity and accuracy throughout the pipeline using tools like Airflow and dbt to manage complex data workflows.

Does this data engineering skill work with Kafka for real-time data processing?

Yes, this data engineering skill works with Kafka for real-time data processing. It supports distributed computing and real-time pipeline construction alongside Python, SQL, Spark, and Airflow to handle streaming data workflows.

What's the best way to optimize ETL/ELT workflows and ensure data quality?

Optimize ETL/ELT workflows and ensure data quality by implementing DataOps patterns, performance tuning, and data validation checks. This Skill enhances transformation efficiency while validating data integrity at each pipeline step using dbt and Airflow.

Can I use this to design data architecture and implement data governance?

Yes, you can use this to design data architecture and implement data governance. It covers data modeling, pipeline orchestration, security, and cost optimization to establish robust data infrastructure for business intelligence.