llmbench

Run a prompt across four models in parallel and render a 2x2 comparison video.

1|Updated Feb 14, 2026
One-click install
npx skills add https://github.com/az9713/my-agent-skills --skill llmbench
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: llmbench
Source: https://github.com/az9713/my-agent-skills/tree/main/.claude/skills/llmbench
Command: npx skills add https://github.com/az9713/my-agent-skills --skill llmbench

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill eliminates the guesswork in comparing large language models by automating a full benchmark pipeline that runs a given prompt across four models in parallel and produces a shared, visual comparison video.

Core Features & Use Cases

  • Parallel model execution across four model endpoints using OpenCode to speed up benchmarking.
  • Remotion-based 2x2 video grid that visually juxtaposes outputs for easy assessment.
  • Use cases include prompt design evaluation, model capability comparison, and end-to-end workflow testing across model families.

Quick Start

Run the llmbench pipeline with a prompt to test four models in parallel and render the comparison video.

Frequently Asked Questions about llmbench

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark multiple LLMs in parallel with a single prompt?

You can benchmark multiple LLMs in parallel by running a single prompt across four model endpoints simultaneously using OpenCode, which automates the execution to speed up the comparison process.

Can I render a side-by-side video comparison of LLM outputs?

Yes, you can render a side-by-side video comparison of LLM outputs using a Remotion-based 2x2 video grid that visually juxtaposes the results for easy assessment.

What is the best way to compare model capabilities across different families?

The best way to compare model capabilities across different families is using an automated benchmark pipeline that executes your prompt across four models in parallel and renders a visual comparison video.

Do I need Remotion to visualize prompt design evaluation results?

Yes, you need Remotion to visualize prompt design evaluation results because it renders the 2x2 video grid that juxtaposes the parallel outputs for easy assessment.

Does the benchmark workflow support end-to-end orchestration for testing model families?

Yes, the benchmark workflow supports end-to-end orchestration for testing model families by handling parallel task execution, Remotion video rendering, and Bash-based workflow automation.

Why use a 2x2 video grid for LLM evaluation instead of text logs?

You use a 2x2 video grid for LLM evaluation because it visually juxtaposes outputs in a shared format, eliminating the guesswork in comparing model capabilities and prompt design results.