measurement-critic

Evaluate measurement instruments and composite scoring functions for validity, reliability, and bias.

Updated Mar 5, 2026
One-click install
npx skills add https://github.com/zivtech/joyus-desktop --skill measurement-critic
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: measurement-critic
Source: https://github.com/zivtech/joyus-desktop/tree/main/.claude/skills/measurement-critic
Command: npx skills add https://github.com/zivtech/joyus-desktop --skill measurement-critic

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill addresses the risk of invalid, biased, or unreliable measurement instruments and scoring functions that can lead to flawed decision-making and Goodhart's Law failures.

Core Features & Use Cases

  • LLM-instrument mode: Evaluates the construct validity, reliability, and bias of measurement designs and results.
  • Composite-scorer mode: Reviews weighted scoring functions for proxy validity, sensitivity, and adversarial robustness.
  • Use Case: Use this before deploying a new automated scoring system to ensure your metrics are actually measuring what you intend and are resistant to manipulation.

Quick Start

Invoke the measurement-critic skill to review the validity of the current scoring function and its component weights.

Frequently Asked Questions about measurement-critic

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I validate the construct validity and reliability of my LLM-based measurement instruments?

Validating measurement instruments requires evaluating construct validity, reliability, and bias across LLM-based designs. Multi-perspective assessments provide statistical, methodological, and domain-expert critiques to ensure your instruments measure the intended constructs accurately and resist manipulation.

How do I review weighted scoring functions for proxy validity and adversarial robustness?

Reviewing weighted scoring functions for proxy validity and adversarial robustness prevents Goodhart's Law exposure. This composite-scorer mode evaluates sensitivity and manipulation resistance to ensure your hand-tuned weighted ranking systems measure intended metrics accurately.

When should I evaluate my automated scoring system to prevent Goodhart's Law failures?

Evaluate your automated scoring system before deployment to prevent Goodhart's Law failures. Pre-deployment reviews of measurement designs and scoring functions ensure your metrics measure intended constructs accurately and resist manipulation, avoiding flawed decision-making.

What is the best way to check my composite scoring function for bias and measurement validity?

The best way to check composite scoring functions for bias and measurement validity is through multi-perspective assessments. Combining statistical, methodological, and domain-expert critiques evaluates calibration adequacy, proxy validity, and sensitivity to ensure robust measurement designs.

Does this measurement validation approach work for both LLM-based designs and hand-tuned weighted ranking systems?

Yes, measurement validation operates across both LLM-based measurement designs and hand-tuned weighted ranking systems. LLM-instrument mode evaluates construct validity and bias, while composite-scorer mode reviews proxy validity, sensitivity, and adversarial robustness for weighted functions.