senior-data-engineer

Design production-grade data pipelines with Python, Spark, Airflow, dbt, and Kafka.

1|Updated Feb 13, 2025
One-click install
npx skills add https://github.com/Aniket-a14/Wizard-w1 --skill senior-data-engineer-aniket-a14
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: senior-data-engineer
Source: https://github.com/Aniket-a14/Wizard-w1/tree/main/.gemini/skills/senior-data-engineer
Command: npx skills add https://github.com/Aniket-a14/Wizard-w1 --skill senior-data-engineer-aniket-a14

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides senior-level guidance and tooling patterns for building scalable data pipelines, ETL/ELT architectures, governance, and DataOps in production-grade environments. It helps teams design robust data infrastructure and processes from planning to deployment.

Core Features & Use Cases

  • End-to-end data pipeline design: ingestion, transformation, orchestration, and monitoring for large-scale datasets.
  • Data governance and quality patterns: validation, lineage, security, and compliance in data workflows.
  • Production ops guidance: deployment patterns, observability, cost optimization, and team leadership guidance.

Quick Start

Design a production-grade data pipeline for ingesting 1TB of daily logs, orchestrated with Airflow, processed with Spark, validated with data quality checks, and monitored with Prometheus or similar observability tools.

Frequently Asked Questions about senior-data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design scalable data pipelines for production environments?

Design scalable data pipelines by structuring ingestion, transformation, and orchestration across modern data stacks like Airflow and Spark. This approach ensures robust data infrastructure, operational excellence, and optimized performance for large-scale datasets.

What is the best way to orchestrate ETL workflows with Airflow and Spark?

Orchestrate ETL workflows by integrating Airflow for scheduling and Spark for distributed data processing. This combination provides production-grade operational excellence, enabling reliable transformation and monitoring for large-scale data pipelines.

How do I implement data observability and governance in ELT pipelines?

Implement data observability and governance by applying validation, lineage tracking, and security compliance patterns within your ELT workflows. This ensures high data quality and operational monitoring across the entire data infrastructure.

Can I use dbt for transformations within a Spark-based data infrastructure?

Yes, dbt can be used for transformations within a Spark-based data infrastructure to manage ELT processes. This integration allows teams to apply structured transformation patterns, governance, and data quality checks seamlessly.

How do I monitor and optimize costs for large-scale data pipelines?

Monitor and optimize data pipeline costs by implementing production ops guidance, observability tools, and performance optimization patterns. This reduces infrastructure overhead while maintaining robust data processing and operational excellence.