run-experiment

Design and execute commit-bound natural experiments on live message corpora.

1|Updated Feb 19, 2026
One-click install
npx skills add https://github.com/mrrts/WorldThreads --skill run-experiment-mrrts
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: run-experiment
Source: https://github.com/mrrts/WorldThreads/tree/main/.agents/skills/run-experiment
Command: npx skills add https://github.com/mrrts/WorldThreads --skill run-experiment-mrrts

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Design and execute rigorous commit-bound natural experiments against a live message corpus to validate hypotheses and measure the impact of prompt-stack changes.

Core Features & Use Cases

  • Structured hypothesis audition with 2–3 candidate hypotheses
  • Pre-registered predictions and explicit success/failure criteria
  • End-to-end experiment design, execution, and reporting with traceability to commits
  • Per-character or per-group scope evaluation and registry integration for reproducibility
  • Documentation-driven workflow that links to reports and project rubrics

Quick Start

Outline a hypothesis, choose a commit boundary, and run a 12-message evaluate with a pre-registered rubric.

Frequently Asked Questions about run-experiment

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I run natural experiments on a live message corpus?

Natural experiments on a live message corpus are run by defining 2–3 candidate hypotheses, pre-registering a prediction with explicit success criteria, and executing against commit boundaries with a structured rubric for honest interpretation.

What is a commit-bound experiment for prompt-stack validation?

A commit-bound experiment validates prompt-stack changes by anchoring natural experiment execution to specific commit boundaries in a live message corpus, ensuring traceability and reproducibility of the measured impact.

How do I pre-register a prediction for data analysis experiments?

Pre-registering a prediction involves defining explicit success and failure criteria alongside a structured rubric before execution, then recording the honest interpretation of results in a registered report for full reproducibility.

Can I evaluate prompt changes with a structured rubric on live data?

Yes, evaluating prompt changes on live data uses a structured rubric across per-character or per-group scope boundaries, executing a defined message evaluation to measure impact and record findings in a registry.

What's the best way to audition hypotheses before running an experiment?

Auditioning hypotheses involves selecting 2–3 candidate hypotheses, choosing a commit boundary, and running a structured evaluation with a pre-registered rubric to ensure rigorous and traceable experiment design.