serving-llms-vllm
Configure vLLM for high-throughput, low-latency LLM serving with PagedAttention and tensor parallelism.
npx skills add https://github.com/KarlinskyS/hermesSkills --skill serving-llms-vllm-karlinskys
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill. Skill: serving-llms-vllm Source: https://github.com/KarlinskyS/hermesSkills/tree/main/mlops/inference/vllm Command: npx skills add https://github.com/KarlinskyS/hermesSkills --skill serving-llms-vllm-karlinskys