dgx-spark-vllm

Optimize vLLM inference deployment on DGX Spark clusters.

7|Updated Dec 25, 2025
One-click install
npx skills add https://github.com/osoleve/the-fold --skill dgx-spark-vllm
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: dgx-spark-vllm
Source: https://github.com/osoleve/the-fold/tree/main/.claude/skills/dgx-spark-vllm
Command: npx skills add https://github.com/osoleve/the-fold --skill dgx-spark-vllm

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill provides structured guidance for deploying and optimizing vLLM on NVIDIA DGX Spark, reducing setup friction and improving inference throughput.

Core Features & Use Cases

  • Deployment Guidance: Best practices for configuring vLLM on DGX Spark clusters, including model selection, server settings, and workload management.
  • Performance Tuning: Recommendations for context size, batching, memory usage, and parallelism to maximize throughput and minimize latency.
  • Use Case: Run chat, coding assistance, and reasoning workloads on DGX Spark with optimized vLLM settings and real-time monitoring.

Quick Start

Configure the vLLM server on a DGX Spark cluster and run a basic chat prompt to validate setup and performance.

Frequently Asked Questions about dgx-spark-vllm

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I optimize LLM inference throughput on DGX Spark using vLLM?

To optimize LLM inference throughput on DGX Spark using vLLM, you tune context size, batching, memory usage, and parallelism settings. This Skill provides structured guidance for configuring these server parameters to maximize performance.

What is the best way to configure a vLLM server for chat and reasoning workloads on DGX Spark?

The best way to configure a vLLM server for chat and reasoning workloads on DGX Spark is by applying recommended deployment settings for model selection, server configuration, and workload management. This Skill outlines the best practices for this setup.

Do I need CUDA-enabled infrastructure to run HuggingFace models with vLLM on DGX Spark?

Yes, you need CUDA-enabled infrastructure to run HuggingFace models with vLLM on DGX Spark. A running vLLM server and compatible CUDA setup are required dependencies for deploying and optimizing your inference workloads.

How do I reduce setup friction when deploying vLLM on DGX Spark clusters?

You can reduce setup friction when deploying vLLM on DGX Spark clusters by following structured deployment guidance. This includes best practices for model selection, server settings, and workload management to streamline the configuration process.

What performance tuning options are available for vLLM inference optimization on DGX Spark?

Performance tuning options for vLLM inference optimization on DGX Spark include adjusting context size, batching, memory usage, and parallelism. These settings maximize throughput and minimize latency for high-throughput AI workloads.