benchmark-agents

Execute AI agent benchmarks across Vercel platform features and the full eval loop.

Updated Dec 12, 2025
One-click install
npx skills add https://github.com/global-os/gproxy --skill benchmark-agents-global-os
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: benchmark-agents
Source: https://github.com/global-os/gproxy/tree/main/.agents/skills/benchmark-agents
Command: npx skills add https://github.com/global-os/gproxy --skill benchmark-agents-global-os

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires vercel-plugin, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides advanced AI agent benchmark scenarios designed to test and push the limits of Vercel's cutting-edge platform features for complex, multi-system builds.

Core Features & Use Cases

  • AI Agent Benchmarking: Pushes Vercel's platform features like Workflow DevKit, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration.
  • Stress-Testing: Designed to stress-test skill injection for complex, multi-system builds.
  • Eval Loop Coverage: Covers the full eval loop: setup → launch → monitor → verify → fix → release → repeat.
  • Real Claude Code Sessions: Launch real Claude Code sessions with the plugin installed, verify skill injection, monitor PostToolUse validation catches, and produce a coverage report.

Quick Start

Launch the benchmark-agents skill by executing the following command in your terminal: benchmark-agents launch

Frequently Asked Questions about benchmark-agents

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark AI agent performance on Vercel platform features?

You can benchmark AI agent performance on Vercel platform features by launching real Claude Code sessions with the vercel-plugin installed. The benchmark evaluates the full eval loop, including setup, launch, monitoring, verification, fixing, and release.

What Vercel platform features can I stress-test for complex multi-system builds?

You can stress-test Vercel platform features including Workflow DevKit, AI Gateway, MCP, Chat SDK, Queues, Flags, Sandbox, and multi-agent orchestration. This evaluates skill injection and validation catches for complex builds.

Does AI agent benchmarking cover the full evaluation loop from setup to release?

AI agent benchmarking covers the full evaluation loop from setup to release. It launches real Claude Code sessions, verifies skill injection, monitors PostToolUse validation catches, and produces a coverage report.

How do I launch AI agent benchmarking sessions in my terminal?

To launch AI agent benchmarking sessions in your terminal, execute the command `benchmark-agents launch`. This initiates real Claude Code sessions with the vercel-plugin installed to monitor and verify platform features.

Do I need the vercel-plugin to run multi-agent orchestration benchmarks?

You need the vercel-plugin dependency to run multi-agent orchestration benchmarks. The benchmark launches real Claude Code sessions with the plugin installed to verify skill injection and monitor PostToolUse validation catches.