sglang-auto-benchmark

Automate repeatable LLM benchmarking across hardware and model configurations.

12|2|Updated Mar 22, 2026
One-click install
npx skills add https://github.com/scottgl9/sglang-spark-gb10-optimizations --skill sglang-auto-benchmark-scottgl9
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sglang-auto-benchmark
Source: https://github.com/scottgl9/sglang-spark-gb10-optimizations/tree/main/.claude/skills/sglang-auto-benchmark
Command: npx skills add https://github.com/scottgl9/sglang-spark-gb10-optimizations --skill sglang-auto-benchmark-scottgl9

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill simplifies and accelerates the process of tuning large language models for optimal performance and cost-efficiency.

Core Features & Use Cases

  • Automated Benchmarking: Executes repeatable, AI-driven performance searches across models and configurations.
  • Data Preparation & Validation: Handles dataset conversion, validation, and scenario setup customized to user needs.
  • Use Case: For a model deployment team, run automated benchmarks on different hardware setups to identify the best configuration for desired throughput and latency targets.

Quick Start

Use the skill to run a benchmarking workflow on your model by providing your dataset and configuration details as commanded by the AI assistant.

Frequently Asked Questions about sglang-auto-benchmark

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate benchmarking for large language models across different hardware?

Automate benchmarking for large language models by executing repeatable, AI-driven performance searches across diverse hardware and model configurations. This process ensures comprehensive coverage of performance metrics like QPS, latency, and throughput.

Can I customize dataset preparation for LLM performance testing?

Customize dataset preparation for LLM performance testing by utilizing automated dataset conversion, validation, and scenario setup tailored to specific user needs and hardware environments.

What is the best way to measure QPS and latency during model optimization?

Measure QPS and latency during model optimization by running automated benchmarking workflows that ensure comprehensive coverage of performance metrics, supporting resume and logging features for reliable tracking.

Does automated benchmarking support resume and logging features for long searches?

Automated benchmarking supports resume and logging features to handle long AI-driven performance searches. This ensures repeatable execution and tracking of performance metrics across various model configurations.

How do I tune AI performance targets for cost-efficiency and throughput?

Tune AI performance targets for cost-efficiency by deploying automated benchmarking workflows that execute repeatable searches. This identifies optimal configurations for throughput, latency, and QPS metrics across diverse hardware.