What problem does it solve?
This Skill automates the creation and maintenance of robust, scalable data pipelines and lakehouse architectures, transforming raw data into trusted, analytics-ready assets.
Core Features & Use Cases
- ETL/ELT Pipeline Development: Design and build idempotent, observable, and self-healing data pipelines.
- Lakehouse Architecture: Implement Medallion Architecture (Bronze, Silver, Gold) on cloud platforms.
- Data Quality & Reliability: Enforce data contracts, monitor SLAs, and implement lineage tracking.
- Streaming Data: Build event-driven pipelines with Kafka and stream processing frameworks.
- Use Case: Automatically ingest data from multiple sources, cleanse and conform it in the Silver layer, and aggregate it into business-ready metrics in the Gold layer, ensuring data quality and timely delivery.
Quick Start
Use the data-eng skill to build a bronze layer pipeline for ingesting JSON data from '/path/to/source' into 's3://my-bucket/bronze/events'.