data-engineer

Design batch and streaming data pipelines with governance and quality checks.

Updated Mar 17, 2026
One-click install
npx skills add https://github.com/involvex/tt2-build-wizard --skill data-engineer-involvex
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-engineer
Source: https://github.com/involvex/tt2-build-wizard/tree/main/.gemini/skills/data-engineer
Command: npx skills add https://github.com/involvex/tt2-build-wizard --skill data-engineer-involvex

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Solves the challenge of building robust, scalable data pipelines and modern analytics platforms that support batch and streaming workloads, data warehousing, and governance across complex environments.

Core Features & Use Cases

  • End-to-end pipeline design for batch and streaming data.
  • Data warehousing and lakehouse architectures with governance and quality checks.
  • Orchestration and automation using modern tooling to ensure reliability and reproducibility.

Quick Start

Design a scalable data pipeline for a provided dataset and production environment.

Frequently Asked Questions about data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
What is the best way to build scalable data pipelines for streaming and batch workloads?

Scalable data pipelines support both batch and streaming workloads using end-to-end architecture design, ensuring reliability and reproducibility through modern orchestration and automation tooling across diverse data environments.

How do I design a lakehouse architecture with proper data governance and quality checks?

Lakehouse architectures integrate data warehousing and governance by applying automated quality checks and cost-aware deployment patterns to ensure data reliability across complex analytics platforms.

Can I use this for end-to-end architecture design across cloud-native data services?

Yes, end-to-end architecture design applies to cloud-native data services, covering everything from batch and streaming pipelines to data warehousing and governance across diverse production environments.

How do I orchestrate ETL pipelines to ensure reproducibility and data quality?

Orchestrate ETL pipelines using modern tooling for automation, which enforces data quality checks and governance policies to guarantee reliable, repeatable deployment patterns for your data warehouse or lakehouse.

When do I need to implement data governance in my analytics platform?

Implement data governance in your analytics platform when building robust lakehouse architectures or scalable data pipelines that require strict data quality, cost awareness, and reliable orchestration across complex environments.