zicato-design-experiment

Author pre-run hypotheses for zicato experiments from observed loss patterns.

4|2|Updated May 14, 2026
One-click install
npx skills add https://github.com/pedapudi/zicato --skill zicato-design-experiment
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: zicato-design-experiment
Source: https://github.com/pedapudi/zicato/tree/main/skills/zicato-design-experiment
Command: npx skills add https://github.com/pedapudi/zicato --skill zicato-design-experiment

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps operators design the next zicato experiment without inventing predictions after seeing the results, connecting observed loss patterns to an allowed mutation target and a testable expected outcome.

Core Features & Use Cases

  • Loss-Grounded Reasoning: Examine telemetry, epoch goals, drift metrics, and outcome failure-mode profiles before choosing what to change.
  • Mutation Validation: Select only permitted mutation points and justify how they could affect the observed failure pattern.
  • Structured Hypothesis Authoring: Produce the required hypothesis schema with predicted drift movements, metric movements, pass-rate expectations, and risks.
  • Use Case: When confabulation risk is concentrated in research tasks, use this Skill to formulate a pre-run hypothesis that tightens the researcher instructions and predicts both quality improvements and possible tool-use costs.

Quick Start

Ask the zicato design-experiment skill to review the current epoch's loss patterns and draft a schema-valid pre-run hypothesis for the next experiment.

Frequently Asked Questions about zicato-design-experiment

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I design a pre-run hypothesis for multi-agent system experiments?

To design a pre-run hypothesis for multi-agent system experiments, analyze telemetry, drift metrics, and failure-mode profiles to connect observed loss patterns to permitted harness mutations. This prevents inventing predictions after seeing results by requiring read-only grounding and structured metric predictions.

Why does experiment design require validating mutation points before running changes?

Experiment design requires validating mutation points to ensure operators select only permitted harness targets and justify how changes affect observed failure patterns. This validation enforces strict separation between predictions and post-run outcomes, preventing confabulation and maintaining evidence-based reasoning.

What's the best way to structure experiment hypotheses using telemetry analysis?

The best way to structure experiment hypotheses using telemetry analysis is producing a schema-valid hypothesis with predicted drift movements, metric movements, pass-rate expectations, and risk analysis. Ground these predictions in read-only telemetry data and epoch goals before proposing changes.

Can I use epoch analysis to predict pass-rate expectations for mutation planning?

Yes, you can use epoch analysis to predict pass-rate expectations for mutation planning by examining outcome failure-mode profiles and drift metrics. This grounds your predictions in observed loss patterns and validates them against allowed mutation points before running the experiment.

Does evidence-based hypothesis authoring work for research task confabulation risk?

Evidence-based hypothesis authoring works for research task confabulation risk by formulating pre-run hypotheses that tighten researcher instructions and predict both quality improvements and possible tool-use costs. This approach requires read-only grounding and structured drift predictions to maintain integrity.

When should I not use pre-run hypothesis authoring for experiment design?

You should not use pre-run hypothesis authoring when telemetry data is unavailable or when mutation points cannot be validated against permitted harness targets. The approach requires read-only grounding, structured drift metrics, and strict separation between predictions and post-run outcomes to function correctly.