data-engineer

Designs scalable data pipelines and manages them across hybrid IT and OT environments.

Updated Aug 20, 2025
One-click install
npx skills add https://github.com/robertlupo1997/open-vocabulary-object-detection --skill data-engineer-robertlupo1997
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-engineer
Source: https://github.com/robertlupo1997/open-vocabulary-object-detection/tree/main/.claude/skills/data-engineer
Command: npx skills add https://github.com/robertlupo1997/open-vocabulary-object-detection --skill data-engineer-robertlupo1997

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Data engineers design and operate scalable data pipelines and infrastructure to manage growing data volumes, ensure reliable ETL/ELT processing, and optimize cost and performance across platforms.

Core Features & Use Cases

  • Design and implement scalable data pipelines using Spark, Kafka, Flink, or Beam, supporting ETL and ELT patterns.
  • Optimize queries and data infrastructure across Snowflake, BigQuery, and Redshift for cost and performance.
  • Orchestrate pipelines with Airflow, Prefect, or Dagster and enforce data quality and governance.
  • Enable lakehouse architectures and data quality frameworks to maintain reliable, auditable data workflows.

Quick Start

Describe your data workflow goals and I will help you design and implement scalable pipelines and infrastructure.

Frequently Asked Questions about data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design scalable data pipelines for ETL and ELT workflows?

To design scalable data pipelines for ETL and ELT, structure processing logic across frameworks like Spark, Kafka, Flink, or Beam. This Skill provides architectural patterns and robust guardrails to build reliable data infrastructure for analytics workloads.

What is the best way to optimize data warehouse costs across Snowflake, BigQuery, and Redshift?

Optimizing data warehouse costs across Snowflake, BigQuery, and Redshift requires applying query tuning and infrastructure governance. This Skill implements cost governance best practices to reduce compute spend and improve performance.

How does data orchestration work with Airflow, Prefect, or Dagster?

Data orchestration with Airflow, Prefect, or Dagster coordinates automated pipeline execution, dependency management, and data quality enforcement. This Skill helps configure these tools to maintain reliable, auditable data workflows.

Can I use this to build a lakehouse architecture with data quality frameworks?

Yes, you can use this Skill to build lakehouse architectures integrated with data quality frameworks. It supports designing reliable, auditable data workflows that enforce governance and quality checks across your infrastructure.

When do I need to choose Spark or Flink for my data engineering pipelines?

Choosing Spark or Flink for data engineering pipelines depends on whether your workloads require batch or stream processing. This Skill guides architecture selection to match the appropriate tool with your pipeline patterns and scalability requirements.