databricks-synthetic-data-gen

Generate realistic synthetic data for Databricks using Spark and Faker.

Updated Jun 11, 2026
One-click install
npx skills add https://github.com/Zack2626-ok/DATN_Website-Dat-Ban-Va-Quan-Ly-Nha-Hang --skill databricks-synthetic-data-gen-zack2626-ok
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: databricks-synthetic-data-gen
Source: https://github.com/Zack2626-ok/DATN_Website-Dat-Ban-Va-Quan-Ly-Nha-Hang/tree/main/.windsurf/skills/databricks-synthetic-data-gen
Command: npx skills add https://github.com/Zack2626-ok/DATN_Website-Dat-Ban-Va-Quan-Ly-Nha-Hang --skill databricks-synthetic-data-gen-zack2626-ok

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires databricks-connect, faker, numpy, pandas, holidays, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill addresses the challenge of creating realistic synthetic data for Databricks, enabling users to simulate real-world scenarios without exposing sensitive information.

Core Features & Use Cases

  • Realistic Data Generation: Utilizes Spark + Faker to generate synthetic data that mimics real business scenarios.
  • Scalability: Supports data generation from thousands to millions of rows.
  • Story-Driven Approach: Ensures data tells a business story, with clear patterns and actionable insights.
  • Use Case: For a new product launch, generate synthetic customer data to simulate user behavior and test marketing strategies.

Quick Start

Generate synthetic customer data for the next product launch by using the 'databricks-synthetic-data-gen' skill.

Frequently Asked Questions about databricks-synthetic-data-gen

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate realistic synthetic data for testing in Databricks?

To generate realistic synthetic data for Databricks, you can use Spark combined with Faker to simulate real-world business scenarios. This approach creates scalable, story-driven datasets suitable for testing and analysis without exposing sensitive information.

Can I generate millions of rows of synthetic data using Spark and Faker?

Yes, you can generate synthetic data using Spark and Faker at scale, supporting datasets ranging from thousands to millions of rows. The generation process ensures data tells a business story with clear patterns and actionable insights for large-scale simulations.

Do I need Databricks Connect to run synthetic data generation?

Yes, Databricks Connect is required for serverless execution when generating synthetic data. You also need Spark and Faker installed to handle the core data generation and anonymization processes within the Databricks environment.

What is the best way to anonymize sensitive data for Databricks analysis?

The best way to anonymize sensitive data for Databricks is generating synthetic data that mimics real business scenarios. By using Spark and Faker, you create realistic datasets for training and analysis, ensuring no actual sensitive information is exposed.

How does story-driven synthetic data generation work for product launches?

Story-driven synthetic data generation creates datasets with clear patterns and actionable insights that simulate user behavior. For a product launch, it generates synthetic customer data to test marketing strategies and model realistic user interactions within Databricks.