LongMemEval Judge
Evaluate QA answers in the LongMemEval benchmark using Anthropic or OpenAI services.
npx skills add https://github.com/tmuskal/arc-agi-benchmarker --skill longmemeval-judge
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill. Skill: LongMemEval Judge Source: https://github.com/tmuskal/arc-agi-benchmarker/tree/main/plugins/longmemeval-benchmarker/skills/judge Command: npx skills add https://github.com/tmuskal/arc-agi-benchmarker --skill longmemeval-judge