gpu-bench

Benchmark GPU-accelerated Moon workloads against CPU baselines in CUDA environments.

2|Updated Mar 26, 2026
One-click install
npx skills add https://github.com/pilotspace/moon --skill gpu-bench
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: gpu-bench
Source: https://github.com/pilotspace/moon/tree/main/.claude/skills/gpu-bench
Command: npx skills add https://github.com/pilotspace/moon --skill gpu-bench

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

GPU-accelerated benchmarking of Moon workloads to compare against CPU baselines, guiding optimization decisions.

Core Features & Use Cases

  • Benchmark vector-distance computations across multiple dimensions and configurations.
  • Measure batch throughput for large-scale operations and memory-bound workloads.
  • Verify CUDA-enabled performance on Moon deployments with deterministic results.

Quick Start

Run gpu-bench to benchmark vector distances, batch throughput, and full GPU tests on a CUDA-enabled system.

Frequently Asked Questions about gpu-bench

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark GPU workloads against CPU baselines for Moon deployments?

You can benchmark GPU workloads against CPU baselines for Moon deployments by running gpu-bench to evaluate vector-distance, batch throughput, and full GPU benchmark scenarios with deterministic results.

What GPU benchmark scenarios are available for vector-distance and batch throughput testing?

Available GPU benchmark scenarios for vector-distance and batch throughput testing include measuring large-scale operations, memory-bound workloads, and multi-dimensional configurations across CUDA-enabled environments.

Do I need the CUDA toolkit to run GPU-accelerated benchmarks on Moon?

Yes, you need the CUDA toolkit to run GPU-accelerated benchmarks on Moon, as the process requires CUDA-enabled environments and GPU-enabled builds with the gpu-cuda feature.

Can I measure batch throughput for memory-bound workloads using GPU benchmarking?

Yes, you can measure batch throughput for memory-bound workloads using GPU benchmarking to evaluate large-scale operations and compare performance against CPU baselines.

How do I verify CUDA-enabled performance with deterministic bench scripts?

Verify CUDA-enabled performance with deterministic bench scripts by running gpu-bench on a CUDA-enabled system to test vector distances, batch throughput, and full GPU functionality.