synthetic-data-generation

Design and validate synthetic data benchmarks for mechanistic interpretability experiments.

4|1|Updated May 20, 2026
One-click install
npx skills add https://github.com/concordance-co/xenon --skill synthetic-data-generation-concordance-co
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: synthetic-data-generation
Source: https://github.com/concordance-co/xenon/tree/main/.agents/skills/synthetic-data-generation
Command: npx skills add https://github.com/concordance-co/xenon --skill synthetic-data-generation-concordance-co

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This guide helps researchers design and repair benchmarks by generating controlled synthetic data that preserves the intended decision bottlenecks, enabling robust mechanistic interpretability studies.

Core Features & Use Cases

  • Design experiments that isolate latent variables and reduce surface-level cues in benchmarks.
  • Repair problematic benchmarks by systematically varying prompts, anchors, and aliases to break surface-label correlations.
  • Use this framework to build audit-ready synthetic datasets that support robust analysis of model reasoning and behavior.

Quick Start

Provide a minimal design prompt and run a small pilot test to validate the core latent variable is being measured.

Frequently Asked Questions about synthetic-data-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
What is synthetic data generation for mechanistic interpretability experiments?

Synthetic data generation creates controlled benchmarks that preserve intended decision bottlenecks. It enforces explicit latent-variable design to isolate model reasoning and reduce surface-level cues, enabling robust mechanistic interpretability studies and behavioral sanity validation.

How do I repair a problematic benchmark to break surface-label correlations?

To repair a problematic benchmark, systematically vary prompts, anchors, and aliases to break surface-label correlations. This anti-shortcut crafting framework enforces explicit latent-variable design, ensuring datasets measure true reasoning rather than relying on superficial dataset artifacts.

How do I design experiments that isolate latent variables in synthetic datasets?

Design experiments by providing a minimal design prompt and running a small pilot test to validate the core latent variable. This framework enforces explicit latent-variable design and anti-shortcut crafting, ensuring your synthetic datasets isolate decision bottlenecks and reduce surface-level cues.

Can I use this framework to build audit-ready synthetic datasets for prompt-variant development?

Yes, you can use this framework to build audit-ready synthetic datasets for prompt-variant development. It provides an audit-friendly framework for prompt inventories and evaluation, supporting robust analysis of model reasoning and behavior across experimental designs.

When should I use anti-shortcut crafting in experimental design?

Use anti-shortcut crafting in experimental design when you need to ensure behavioral sanity by breaking surface-label correlations. It is essential for repairing problematic benchmarks and validating that synthetic data measures intended decision bottlenecks rather than superficial dataset artifacts.