running-experiment-matrices

Coordinates execution and analysis of experiment matrices for grammar-constrained symbolic diffusion models.

1|Updated Jul 12, 2026
One-click install
npx skills add https://github.com/Tyler-R-Kendrick/slm-training --skill running-experiment-matrices
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: running-experiment-matrices
Source: https://github.com/Tyler-R-Kendrick/slm-training/tree/main/.agents/skills/running-experiment-matrices
Command: npx skills add https://github.com/Tyler-R-Kendrick/slm-training --skill running-experiment-matrices

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

This Skill manages the complexity of running, extending, and interpreting multi-dimensional experiment matrices for grammar-constrained symbolic diffusion models, ensuring results are grounded in evidence and architectural invariants.

Core Features & Use Cases

  • Matrix Execution: Provides a standardized interface to run quality, grammar, performance, and phase-based experiment matrices.
  • Result Validation: Enforces strict interpretation rules, including version stamp verification and size-matched arm comparisons, to prevent invalid ship claims.
  • Use Case: When testing a new model lever, use this Skill to run the quality matrix subset, compare results against baseline controls, and generate the required JSON scoreboards for research documentation.

Quick Start

Execute the quality matrix subset for experiment E53 by running the quality matrix script with the specified device and context parameters.

Frequently Asked Questions about running-experiment-matrices

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I run and validate multi-dimensional experiment matrices for symbolic diffusion models?

You run and validate symbolic diffusion model experiment matrices by executing standardized scripts for quality, grammar, performance, and phase-based tests, then enforcing strict size-matched arm comparisons to prevent invalid ship claims.

What is the best way to benchmark grammar-constrained diffusion models across different model levers?

Benchmarking grammar-constrained diffusion models requires running the quality matrix subset to compare new model levers against baseline controls, generating version-stamped JSON scoreboards to ensure empirical rigor in research documentation.

Can I use this execution framework to test performance metrics for small language models (slm)?

Yes, you can use this framework to test small language model performance metrics by executing the specified performance matrix scripts with defined device and context parameters to capture benchmark results.

Why do my experiment matrix results show invalid ship claims during validation?

Invalid ship claims occur when validation rules are violated, specifically failing to verify version stamps or neglecting size-matched arm comparisons against baseline controls during the experiment suite analysis.

How do I enforce architectural invariants when testing new model levers in an experimentation suite?

Enforce architectural invariants by utilizing version-controlled evaluation protocols within the experiment matrix scripts, ensuring all model lever validations are strictly grounded in evidence and structural compliance.

When do I need to use a multi-dimensional experiment matrix for symbolic diffusion validation?

You need a multi-dimensional experiment matrix when validating complex symbolic diffusion models to manage the complexity of running, extending, and interpreting quality, grammar, and performance tests across defined experiment suites.