Agent Skills by Tal Muskal
Showing 13 vetted skills indexed across 1 GitHub repositories.
longmemeval-resume
Resume incomplete LongMemEval benchmark runs from the last checkpoint.
setup
Set up the LongMemEval benchmarking environment with conda and HuggingFace datasets.
cross-harness
Automate LongMemEval benchmarking across Codex, Gemini, and OpenCode harnesses.
compare-runs
Compare two LongMemEval runs by scorecards, harness configurations, and per-question-type accuracy.
LongMemEval Judge
Evaluate QA answers in the LongMemEval benchmark using Anthropic or OpenAI services.
report
Generate LongMemEval benchmark reports in markdown, JSON, or summary formats.
browse-tests
Browse and preview LongMemEval dataset items by question ID.
run-benchmark
Automate LongMemEval benchmark execution with hypothesis generation and scoring.
arc-agi-benchmarker-setup
Install Python, create a virtual environment, and verify the ARC-AGI package setup.
arc-cross-harness
Generate cross-harness benchmarking instructions for ARC-AGI and compare results.
arc-agi-benchmarker:report
Generate detailed ARC-AGI benchmark reports with score breakdowns and completion metrics.
arc-agi-browse-tests
Browse ARC-AGI environments with game details and historical scores.
benchmark-adder
Automate Claude Code plugin creation for benchmarking setups from a repository URL.