evaluation-running
Automate running and back-testing AI safety evaluations across model generations.
npx skills add https://github.com/DouwMarx/evaluating-evaluations --skill evaluation-running-douwmarx
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill. Skill: evaluation-running Source: https://github.com/DouwMarx/evaluating-evaluations/tree/main/evaluation-running Command: npx skills add https://github.com/DouwMarx/evaluating-evaluations --skill evaluation-running-douwmarx