evaluation-running
Run and back-test AI safety evaluations across model generations.
npx skills add https://github.com/EquiStamp/evaluating-evaluations --skill evaluation-running
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill. Skill: evaluation-running Source: https://github.com/EquiStamp/evaluating-evaluations/tree/main/.claude/skills/evaluation-running Command: npx skills add https://github.com/EquiStamp/evaluating-evaluations --skill evaluation-running