experiment-runbook-discipline

Plan long-running experiments with observable artifacts and auditable trails.

Updated Mar 17, 2026
One-click install
npx skills add https://github.com/balandongiv/agent-skillbook --skill experiment-runbook-discipline
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: experiment-runbook-discipline
Source: https://github.com/balandongiv/agent-skillbook/tree/main/skills/experiment-runbook-discipline/exports/claude
Command: npx skills add https://github.com/balandongiv/agent-skillbook --skill experiment-runbook-discipline

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Plan, launch, monitor, and document long-running experiments or validation sweeps with smallest-real-data smoke scopes, fresh experiment prefixes, live status artifacts, rolling logs, and promotion to full runs only after explicit pass criteria are met.

Core Features & Use Cases

  • Standardize experiment workflows with observability and durable artifacts.
  • Provide fresh prefixes and explicit criteria to promote runs to full experiments.
  • Ensure live status artifacts, rolling logs, and audit trails for easier debugging and compliance.

Quick Start

Observe an ongoing experiment and capture its run prefix, dataset scope, and status artifacts.

Frequently Asked Questions about experiment-runbook-discipline

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I monitor long-running experiments and maintain an auditable trail?

You can monitor long-running experiments by generating live status artifacts and rolling logs. This approach ensures durable observability and creates an auditable trail for debugging and compliance throughout the experiment lifecycle.

What is the best way to enforce promotion criteria before scaling validation sweeps to full runs?

Enforce promotion criteria by applying explicit pass conditions before scaling validation sweeps to full runs. This requires generating fresh experiment prefixes and capturing smallest-real-data smoke scopes to verify readiness.

How do I plan pipeline experiments with observable artifacts and run-visibility?

Plan pipeline experiments by standardizing workflows with observable artifacts and run-visibility. This involves tracking status artifacts and enforcing explicit progression criteria to maintain durable experiment governance.

Does experiment governance require fresh prefixes for real-data validation workflows?

Yes, experiment governance requires fresh prefixes for real-data validation workflows to avoid collisions. Fresh prefixes ensure that status tracking and audit trails remain isolated and clearly identifiable during progression.

When do I need to capture smallest-real-data smoke scopes for pipeline experiments?

Capture smallest-real-data smoke scopes for pipeline experiments before promoting validation sweeps to full runs. This ensures that explicit pass criteria are met using minimal data scope before committing to full execution.