test-critic

Audit evaluation suites for benchmark validity and statistical rigor.

Updated Mar 5, 2026
One-click install
npx skills add https://github.com/zivtech/joyus-desktop --skill test-critic-zivtech
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: test-critic
Source: https://github.com/zivtech/joyus-desktop/tree/main/.claude/skills/test-critic
Command: npx skills add https://github.com/zivtech/joyus-desktop --skill test-critic-zivtech

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill prevents the common pitfall of making high-stakes decisions based on flawed or biased evaluation benchmarks, ensuring your testing infrastructure is scientifically sound.

Core Features & Use Cases

  • Evaluation Audit: Performs a multi-perspective review of test fixtures, rubrics, and statistical design to identify contamination, strawman baselines, and overfitting.
  • Evidence-Based Feedback: Provides actionable, evidence-driven findings that help you refine your benchmarks before running large-scale tests.
  • Use Case: Before running a 100-fixture benchmark to compare a new AI skill against a baseline, use this Skill to verify that your rubric isn't teaching-to-the-test and that your sample size is statistically significant.

Quick Start

Use the test-critic skill to review the evaluation design in the current directory for rigor and fairness.

Frequently Asked Questions about test-critic

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I audit an evaluation suite for benchmark validity before running tests?

Auditing an evaluation suite for benchmark validity involves reviewing fixture coverage, rubric fairness, and baseline representativeness to identify design flaws. This process checks for contamination, strawman baselines, and inadequate sample sizes to prevent unreliable performance claims.

What is a strawman baseline in benchmark testing and how do I identify it?

A strawman baseline in benchmark testing is a weak comparison point that makes a new system appear artificially superior. You identify it through an evidence-driven audit that evaluates baseline representativeness and statistical rigor to ensure fair comparisons.

How do I check if my test fixtures have adequate coverage and prevent teaching-to-the-test?

Checking test fixture coverage requires a multi-perspective review of your evaluation design to ensure comprehensive data representation. This prevents teaching-to-the-test by verifying rubric fairness and identifying systemic design flaws like overfitting before large-scale testing.

Can I use statistical rigor checks to verify sample size significance for data science benchmarks?

Statistical rigor checks can verify sample size significance for data science benchmarks by evaluating reproducibility across scientific dimensions. This ensures your evaluation design prevents high-stakes decisions from being based on inadequate sample sizes or biased data.

What are common limitations when evaluating rubric fairness in AI skills?

Limitations when evaluating rubric fairness include systemic design flaws like contamination and overfitting that can skew results. Addressing these requires comprehensive, evidence-driven audits to ensure reproducibility and baseline representativeness across pragmatic dimensions.

Why does my benchmark show high performance but fail in production scenarios?

Benchmarks showing high performance but failing in production often suffer from contamination, inadequate sample sizes, or overfitting rubrics. Performing a rigorous quality gate audit identifies these statistical validity issues before unreliable performance claims are made.