evaluation-building

Construct AI safety evaluation instruments using EquiStamp pipeline and Inspect framework.

Updated Feb 24, 2026
One-click install
npx skills add https://github.com/DouwMarx/evaluating-evaluations --skill evaluation-building-douwmarx
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: evaluation-building
Source: https://github.com/DouwMarx/evaluating-evaluations/tree/main/evaluation-building
Command: npx skills add https://github.com/DouwMarx/evaluating-evaluations --skill evaluation-building-douwmarx

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires inspect-ai, Docker, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill simplifies the process of building AI safety evaluations from validated designs, reducing the complexity and time required for implementation.

Core Features & Use Cases

  • Build AI Safety Evaluations: Create comprehensive evaluation instruments from validated designs.
  • Use Cases: Ideal for implementing evaluations, setting up Inspect tasks, building scoring pipelines, and executing Phase 4 of the EquiStamp eval pipeline.
  • Framework Support: Supports the UK AISI's Inspect framework for building and executing evaluations.

Quick Start

Use the 'evaluation-building' skill to build an evaluation instrument for [capability], following the prompts to gather context and create a build plan document.

Frequently Asked Questions about evaluation-building

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I build an AI safety evaluation from a validated design?

You can build an AI safety evaluation from a validated design by using the EquiStamp eval pipeline to construct evaluation instruments, generating a build plan document through guided prompts to gather context.

Can I use the Inspect framework to set up AI evaluation tasks?

Yes, the Inspect framework is fully supported for building and executing AI safety evaluations, allowing you to set up Inspect tasks and build comprehensive scoring pipelines for your instruments.

Do I need Docker and Python to run Inspect AI evaluation pipelines?

Yes, you need Python and the Inspect AI library for execution, along with Docker for containerized environments to properly run the evaluation building pipeline.

What is the best way to implement Phase 4 of the EquiStamp eval pipeline?

The best way to implement Phase 4 of the EquiStamp eval pipeline is to construct evaluation instruments from validated designs using the provided skill, which simplifies the complexity of setting up scoring pipelines.

How does constructing evaluation instruments reduce implementation time?

Constructing evaluation instruments reduces implementation time by simplifying the process of building AI safety evaluations from validated designs, removing the manual complexity of configuring Inspect tasks.