Tal Muskal avatar

Tal Muskal

Community

@tmuskal

102Followers
|
47Public Repos
|
13Published Skills

Agent Skills by Tal Muskal

Showing 13 vetted skills indexed across 1 GitHub repositories.

tmuskaltmuskal
3

longmemeval-resume

Resume incomplete LongMemEval benchmark runs from the last checkpoint.

Community
Basic
tmuskaltmuskal
3

setup

Set up the LongMemEval benchmarking environment with conda and HuggingFace datasets.

Community
Intermediate
tmuskaltmuskal
3

cross-harness

Automate LongMemEval benchmarking across Codex, Gemini, and OpenCode harnesses.

Community
Intermediate
tmuskaltmuskal
3

compare-runs

Compare two LongMemEval runs by scorecards, harness configurations, and per-question-type accuracy.

Community
Intermediate
tmuskaltmuskal
3

LongMemEval Judge

Evaluate QA answers in the LongMemEval benchmark using Anthropic or OpenAI services.

Community
Intermediate
tmuskaltmuskal
3

report

Generate LongMemEval benchmark reports in markdown, JSON, or summary formats.

Community
Intermediate
tmuskaltmuskal
3

browse-tests

Browse and preview LongMemEval dataset items by question ID.

Community
Intermediate
tmuskaltmuskal
3

run-benchmark

Automate LongMemEval benchmark execution with hypothesis generation and scoring.

Community
Advanced
tmuskaltmuskal
3

arc-agi-benchmarker-setup

Install Python, create a virtual environment, and verify the ARC-AGI package setup.

Community
Advanced
tmuskaltmuskal
3

arc-cross-harness

Generate cross-harness benchmarking instructions for ARC-AGI and compare results.

Community
Intermediate
tmuskaltmuskal
3

arc-agi-benchmarker:report

Generate detailed ARC-AGI benchmark reports with score breakdowns and completion metrics.

Community
Intermediate
tmuskaltmuskal
3

arc-agi-browse-tests

Browse ARC-AGI environments with game details and historical scores.

Community
Advanced
tmuskaltmuskal
3

benchmark-adder

Automate Claude Code plugin creation for benchmarking setups from a repository URL.

Community
Advanced