evaluation-running

Automate running and back-testing AI safety evaluations across model generations.

Updated Feb 24, 2026
One-click install
npx skills add https://github.com/DouwMarx/evaluating-evaluations --skill evaluation-running-douwmarx
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluation-running
Source: https://github.com/DouwMarx/evaluating-evaluations/tree/main/evaluation-running
Command: npx skills add https://github.com/DouwMarx/evaluating-evaluations --skill evaluation-running-douwmarx

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

The evaluation-running skill addresses the challenge of running and back-testing AI safety evaluations across different model generations. It automates the process of validating instruments, executing evals, and analyzing results, enabling efficient and thorough safety assessments.

Core Features & Use Cases

  • Evaluation Automation: Automate the execution of AI safety evaluations across multiple models and generations.
  • Validation of Instruments: Validate the instruments used in evaluations to ensure accuracy and reliability.
  • Analysis of Results: Provide detailed analysis of evaluation results, including statistical tests and trend lines.
  • Use Case: Use this skill to validate an AI safety evaluation instrument across various model generations and analyze the results to identify any safety issues.

Quick Start

Run the evaluation-running skill with the following command: /evaluation-running run-eval --evaluation-id [evaluation_id] --model-panel [model_panel] --budget [budget]

Frequently Asked Questions about evaluation-running

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate AI safety evaluations across multiple model generations?

You can automate AI safety evaluations across model generations by running a command that executes evals against a specified model panel and budget, validating instruments, and analyzing results with statistical tests.

What statistical analysis is included when back-testing AI safety evaluations?

Back-testing AI safety evaluations includes statistical tests and trend analysis to validate instruments and identify safety issues across different model generations.

How do I validate AI safety evaluation instruments before running them?

Validating AI safety evaluation instruments is automated by the skill, which checks accuracy and reliability across the specified model panel before executing the full evaluation.

What do I need to set up before running automated AI safety evaluations?

You need access to multiple model generations and evaluation instruments, plus an evaluation ID, model panel, and budget to execute the automated evaluation run.

Can I run AI safety evaluations with a limited compute budget?

Yes, you can specify a budget parameter when running evaluations to control resource consumption while still automating execution and statistical analysis across model generations.

Why use automated evaluation running instead of manual AI safety testing?

Automating AI safety evaluation running enables efficient back-testing across model generations, validating instruments and generating trend analysis that manual testing cannot easily scale to.