evaluating-llms-harness
Evaluate large language models against standardized benchmarks using lm-evaluation-harness.
npx skills add https://github.com/guccang/blogclaw --skill evaluating-llms-harness-guccang
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill. Skill: evaluating-llms-harness Source: https://github.com/guccang/blogclaw/tree/main/cmd/hermes-agent/vendor/hermes_runtime/skills/mlops/evaluation/lm-evaluation-harness Command: npx skills add https://github.com/guccang/blogclaw --skill evaluating-llms-harness-guccang