Evaluation Analyst

Quantify uncertainty and effect size to classify experiment outcomes as supports, inconclusive, or rejects.

1|Updated Feb 8, 2026
One-click install
npx skills add https://github.com/soheunyi/get-research-done --skill evaluation-analyst
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: Evaluation Analyst
Source: https://github.com/soheunyi/get-research-done/tree/main/skills/grd-evaluation-analyst
Command: npx skills add https://github.com/soheunyi/get-research-done --skill evaluation-analyst

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Converts experimental results into defensible decisions by quantifying uncertainty, effect sizes, and significance, reducing overclaiming and ambiguity in hypothesis conclusions.

Core Features & Use Cases

  • Quantifies uncertainty with clear intervals and confidence levels to accompany metrics.
  • Interprets experiment outcomes as supports, inconclusive, or rejects with explicit linkage to baseline and hypotheses.
  • Applies across multiple variants and metrics, referencing predeclared decision thresholds and the workflow rules for reproducibility.

Quick Start

Provide your experiment outcomes and hypothesis details, and the skill will output a decision with quantified uncertainty and effect size.

Frequently Asked Questions about Evaluation Analyst

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I quantify uncertainty and effect size from experiment outputs?

To quantify uncertainty from experiment outputs, provide your metrics, hypothesis details, and predeclared decision thresholds. The analysis outputs defensible conclusions by computing uncertainty ranges, significance levels, and practical effect sizes to support, reject, or deem results inconclusive.

What is the best way to interpret hypothesis-testing results across multiple variants?

Interpreting hypothesis-testing results across multiple variants requires linking evidence to your baseline and applying predeclared decision thresholds. This process compares metrics across variants to explicitly determine whether the data supports, rejects, or remains inconclusive regarding the hypothesis.

How does statistical significance help reduce overclaiming in decision-making?

Statistical significance reduces overclaiming in decision-making by enforcing explicit linkage between evidence and hypotheses. By applying predeclared thresholds and quantifying uncertainty intervals, it prevents ambiguous conclusions and ensures outcomes strictly reflect the experimental data.

When do I need to calculate practical effect size for my experiments?

You need to calculate practical effect size for experiments when interpreting outcomes across multiple variants to decide between supports, inconclusive, or rejects. It quantifies the magnitude of differences, complementing significance testing to ensure defensible, reproducible decision-making.