databricks-spark-declarative-pipelines

Automates creation, configuration, and management of Databricks Lakeflow Spark Declarative Pipelines.

Updated Mar 23, 2024
One-click install
npx skills add https://github.com/m19c/dotfiles --skill databricks-spark-declarative-pipelines-m19c
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: databricks-spark-declarative-pipelines
Source: https://github.com/m19c/dotfiles/tree/main/claude/.claude/skills/databricks-spark-declarative-pipelines
Command: npx skills add https://github.com/m19c/dotfiles --skill databricks-spark-declarative-pipelines-m19c

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires pyspark.pipelines as dp, databricks, delta, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill automates the creation, configuration, and updating of Databricks Lakeflow Spark Declarative Pipelines (SDP/LDP), streamlining the process of building and maintaining complex data pipelines and Delta Live Tables.

Core Features & Use Cases

  • Automated Pipeline Creation: Quickly set up new SDP projects with predefined configurations.
  • Streamlined Pipeline Management: Simplifies the process of updating and refreshing existing pipelines.
  • Data Ingestion: Handles various data ingestion patterns, including streaming tables, materialized views, CDC, and SCD Type 2.
  • Use Case: Use this Skill to create a new SDP project for a data pipeline, configure it with the necessary components, and automate the data ingestion process from various sources.

Quick Start

Use the databricks-spark-declarative-pipelines skill to create a new SDP project named 'my_pipeline'.

Frequently Asked Questions about databricks-spark-declarative-pipelines

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I create a Databricks Lakeflow Spark Declarative Pipeline?

To create a Databricks Lakeflow Spark Declarative Pipeline, you can use this Skill to automate project setup with predefined configurations. It streamlines building and maintaining complex data pipelines and Delta Live Tables.

How does this approach handle CDC and SCD Type 2 data ingestion?

This approach handles CDC and SCD Type 2 data ingestion by automating the configuration of these patterns within Databricks Lakeflow Spark Declarative Pipelines. It manages data ingestion from various sources efficiently.

Do I need specific dependencies to manage Delta Live Tables with this Skill?

Yes, you need specific dependencies to manage Delta Live Tables with this Skill. It requires pyspark.pipelines as dp, databricks, and delta for creating and managing declarative pipelines.

Can I use this to configure Auto Loader ingestion patterns for streaming data?

Yes, you can use this to configure Auto Loader ingestion patterns for streaming data. The Skill automates data ingestion, handling streaming tables and materialized views within Databricks Lakeflow pipelines.

What is the best way to update existing Databricks Lakeflow pipelines?

The best way to update existing Databricks Lakeflow pipelines is using this Skill to streamline management and refreshing processes. It simplifies updating configurations for your Spark Declarative Pipelines.

Why use declarative pipelines for data engineering workflows in Databricks?

Use declarative pipelines for data engineering workflows in Databricks to effortlessly create and manage complex data processing logic. This approach streamlines building Delta Live Tables and automates ingestion.