evaluation-running

Run and back-test AI safety evaluations across model generations.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/EquiStamp/evaluating-evaluations --skill evaluation-running
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluation-running
Source: https://github.com/EquiStamp/evaluating-evaluations/tree/main/.claude/skills/evaluation-running
Command: npx skills add https://github.com/EquiStamp/evaluating-evaluations --skill evaluation-running

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Run and back-test AI safety evaluations across model generations. Use when executing evals, assembling model panels, validating instruments, or running Phase 5 of the EquiStamp eval pipeline.

Core Features & Use Cases

  • Phase 5 guidance including running built evaluations across a panel of models
  • Build model panels, validate instruments, generate time-series, and reconcile runtime data with design assumptions
  • Provide plan for cost estimation, logging, and run-to-run variance tracking

Quick Start

Load your Phase 4 artifact and follow the Run Plan and Execution Guidance to start a Phase 5 run across your model panel.

Frequently Asked Questions about evaluation-running

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I run AI safety evaluations across a panel of models?

Run AI safety evaluations by loading your Phase 4 artifact and following the Run Plan to execute a Phase 5 run across your model panel. This validates instruments and generates time-series runtime data.

What is back-testing AI evaluations and when do I need it?

Back-testing AI evaluations runs built safety assessments across model generations to validate performance over time. You need it when reconciling runtime data with design assumptions and tracking run-to-run variance.

Do I need a Phase 4 artifact to start running model panel evaluations?

Yes, you need a Phase 4 artifact to start running model panel evaluations. Loading this artifact provides the structured run plans and design assumptions required to execute Phase 5 of the eval pipeline.

Can I track run-to-run variance and estimate costs when executing AI evals?

Yes, you can track run-to-run variance and estimate costs when executing AI evals. The execution guidance provides structured plans for cost estimation, logging, and statistical validation during the run.

How do I validate instruments and generate time-series data for AI evaluations?

Validate instruments and generate time-series data for AI evaluations by assembling a configured model panel and following Phase 5 guidance. This ensures rigorous statistical validation and reconciles runtime data with design assumptions.