principal-data-engineer

Guides enterprise data teams in designing, reviewing, and governing scalable data platforms with Iceberg, dbt, Airflow, Polars, and DuckDB.

Updated Jan 15, 2026
One-click install
npx skills add https://github.com/rory-data/copilot --skill principal-data-engineer
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: principal-data-engineer
Source: https://github.com/rory-data/copilot/tree/main/skills/principal-data-engineer
Command: npx skills add https://github.com/rory-data/copilot --skill principal-data-engineer

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires loguru, and includes scripts (resource) and references (resource) components.

What problem does it solve?

Provides expert guidance to design, review, and govern scalable, reliable data platforms, elevating data programs from initial viability to long-term value and governance.

Core Features & Use Cases

  • Data Platform Architecture: emphasizes the -ilities (scalability, reliability, maintainability, observability) and patterns for architecture reviews, system decoupling, and cost-aware design.
  • Pipeline Engineering Standards: sets Airflow and Python code standards, testing regimes, anti-patterns, and code quality guidance.
  • Data Quality & Observability: promotes data contracts, open lineage, data quality checks, and SLA monitoring.
  • Composable Data Stack: guidance on dlt, Polars, DuckDB, Iceberg, dbt, and Open Data Contract Standards for flexible pipelines.
  • Usage Scenarios: architectural reviews, complex debugging, template creation, and standard-setting.

Quick Start

Summarize a scalable data-platform architecture with Iceberg and Airflow, including data contracts and quality gates.

Frequently Asked Questions about principal-data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design a scalable data platform architecture using Iceberg and Airflow?

To design a scalable data platform with Iceberg and Airflow, you need to enforce data contracts, implement idempotent pipelines, and establish OpenLineage tracking. This Skill provides architectural patterns for decoupling systems and ensuring lakehouse reliability across batch and streaming contexts.

What are the essential data quality and observability standards for an enterprise lakehouse?

Essential lakehouse observability standards require data contracts, OpenLineage integration, and rigorous data quality checks. This Skill specifies SLA monitoring and testing regimes including unit, integration, and data quality tests to maintain platform reliability and maintainability.

Does this guidance support setting up dbt documentation and testing regimes for pipeline engineering?

Yes, this guidance establishes dbt documentation and comprehensive testing regimes for pipeline engineering. It defines Airflow and Python code standards, specifies anti-patterns to avoid, and sets quality gates to ensure your data transformations meet enterprise governance requirements.

Can I use Polars and DuckDB within a composable data stack for complex pipeline engineering?

Yes, you can use Polars and DuckDB within a composable data stack. This Skill provides implementation guidance for flexible pipelines using these tools alongside dlt and Iceberg, ensuring your architecture supports both batch and streaming processing efficiently.

What is the best way to govern data platforms and prevent architectural anti-patterns?

The best way to govern data platforms and prevent anti-patterns is through architectural reviews and strict adherence to engineering standards. This Skill guides teams in system decoupling, cost-aware design, and applying Open Data Contract Standards to elevate long-term platform value.