search-benchmark

Benchmark Switchboard's search tool across opus, sonnet, and haiku models.

15|7|Updated Feb 26, 2026
One-click install
npx skills add https://github.com/daltoniam/switchboard --skill search-benchmark
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: search-benchmark
Source: https://github.com/daltoniam/switchboard/tree/main/.agents/skills/search-benchmark
Command: npx skills add https://github.com/daltoniam/switchboard --skill search-benchmark

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Benchmark Switchboard's search tool across models to measure cross-model search quality for tool discovery.

Core Features & Use Cases

  • Cross-model benchmark across opus, sonnet, and haiku in parallel
  • Generates a comparison table and identifies optimization opportunities
  • Use cases: performance assessment of search suggestions, model-tier differences, and when evaluating Phase 2 tag impact

Quick Start

Run the search-benchmark skill to dispatch identical scenarios to opus, sonnet, and haiku in parallel and generate a comparison report

Frequently Asked Questions about search-benchmark

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark cross-model search quality for tool discovery?

You can benchmark cross-model search quality by dispatching identical scenarios to opus, sonnet, and haiku in parallel. This measures search suggestions equally across model tiers and produces a structured comparison report.

What is cross-model search benchmarking used for?

Cross-model search benchmarking is used to evaluate search tool quality for tool discovery. It identifies performance differences across model tiers and assesses Phase 2 tag impact to reveal optimization opportunities.

Can I compare search results across opus, sonnet, and haiku simultaneously?

Yes, you can compare search results across opus, sonnet, and haiku simultaneously. The benchmark dispatches identical scenarios to all three models in parallel to ensure equal treatment and generate a comparison table.

How do I generate a comparison report for multimodel search results?

You generate a comparison report for multimodel search results by running a benchmark that applies synthetic and live-model scenarios across opus, sonnet, and haiku. It collects the results and outputs a structured comparison table.

Does the search benchmark support synthetic and live-model scenarios?

Yes, the search benchmark supports both synthetic and live-model scenarios. It applies these scenarios equally across opus, sonnet, and haiku to compare results and measure cross-model search quality.

What are the limitations of multimodel search benchmarking?

A limitation of multimodel search benchmarking is that it requires equal treatment across opus, sonnet, and haiku to compare results accurately. It focuses strictly on search tool quality and Phase 2 tag impact rather than general model performance.