run-benchmark
Automate LongMemEval benchmark execution with hypothesis generation and scoring.
npx skills add https://github.com/tmuskal/arc-agi-benchmarker --skill run-benchmark-tmuskal
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill. Skill: run-benchmark Source: https://github.com/tmuskal/arc-agi-benchmarker/tree/main/plugins/longmemeval-benchmarker/skills/run-benchmark Command: npx skills add https://github.com/tmuskal/arc-agi-benchmarker --skill run-benchmark-tmuskal