data-engineer

Design and implement scalable ETL/ELT pipelines for batch and streaming data.

Updated Jan 19, 2023
One-click install
npx skills add https://github.com/claudchereji/VisualVerses --skill data-engineer-claudchereji
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: data-engineer
Source: https://github.com/claudchereji/VisualVerses/tree/main/.opencode/skills/data-engineer
Command: npx skills add https://github.com/claudchereji/VisualVerses --skill data-engineer-claudchereji

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill addresses the complexity of building and maintaining robust, scalable, and cost-efficient data infrastructure and pipelines.

Core Features & Use Cases

  • Data Pipeline Development: Design, implement, and optimize ETL/ELT processes for batch and streaming data.
  • Data Platform Architecture: Architect data lakes, data warehouses, and lakehouses on cloud platforms.
  • Big Data Technologies: Leverage expertise in tools like Spark, Kafka, and Flink for large-scale data processing.
  • Use Case: A company needs to ingest real-time sales data from multiple sources, process it, and make it available for analytics dashboards within minutes. This Skill can design and implement such a streaming pipeline.

Quick Start

Use the data-engineer skill to design a scalable data pipeline for real-time sales data ingestion.

Frequently Asked Questions about data-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design scalable data pipelines for real-time ingestion and analytics?

Scalable data pipelines for real-time ingestion are designed by implementing ETL/ELT processes that leverage big data technologies like Spark, Kafka, and Flink. This ensures data is processed and available for analytics dashboards within minutes.

What is the best way to architect a data lakehouse on cloud platforms?

Architecting a data lakehouse on cloud platforms involves designing scalable infrastructure that combines data lake flexibility with warehouse structure. This approach supports reliable, efficient, and cost-optimized storage for business intelligence.

Can I use this approach for both batch and streaming ETL processes?

Yes, the approach supports both batch and streaming ETL/ELT processes. It leverages big data technologies to handle large-scale data ingestion, transformation, and processing while ensuring high availability and low latency.

How do I optimize big data infrastructure for cost and performance?

Optimizing big data infrastructure for cost and performance requires building reliable, efficient, and cost-optimized data platforms. It focuses on scaling data architecture to handle ingestion, storage, and processing without unnecessary overhead.

When should I use Spark and Kafka for large-scale data processing?

Spark and Kafka are used for large-scale data processing when you need to handle high-throughput data ingestion and real-time transformation. They enable streaming pipelines that process incoming data and make it available for analytics within minutes.