What problem does it solve?
Manually creating data quality specifications is time-consuming, error-prone, and often results in incomplete rules that miss upstream defects or lack traceability to upstream engineering artifacts. This Skill eliminates that manual work by enforcing a structured, gate-driven workflow that produces complete, compliant, and build-ready DQS documents aligned to enterprise data quality standards.
Core Features & Use Cases
- Structured Elicitation Workflow: Uses targeted Q&A to gather all required DQ design decisions (severity thresholds, alert channels, tolerance levels) before generating output, eliminating vague or missing rules.
- Mandatory Upstream Validation Gates: Enforces approval of all upstream STM, DMS, and DRD artifacts and validates source data volumes via read-only database queries to ensure statistical baselines are data-driven.
- Spark-Expectations Integration: Auto-generates compatible SE YAML rule files after DQS validation, with correct rule typing and field semantics for seamless pipeline deployment.
- Use Case: A data quality engineer working on a healthcare medallion pipeline can use this Skill to translate approved STM mappings, DMS schemas, and DRD requirements into a full DQS covering field validations, referential integrity, statistical tests, reconciliation, and SLA monitoring in minutes.
Quick Start
Use the create-dqs skill to generate a complete, compliant data quality specification for your latest approved STM, DMS, and DRD artifacts, including all validation rules, reconciliation checks, and alert frameworks for your medallion data pipeline.