comfyui

Generate images, video, and audio content using ComfyUI.

1|Updated Jun 9, 2026
One-click install
npx skills add https://github.com/aivos-xie/hermes-skills --skill comfyui-aivos-xie
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: comfyui
Source: https://github.com/aivos-xie/hermes-skills/tree/main/creative/comfyui
Command: npx skills add https://github.com/aivos-xie/hermes-skills --skill comfyui-aivos-xie

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires comfy-cli, comfyui, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill allows users to generate images, video, and audio content using ComfyUI, a powerful tool for creative applications.

Core Features & Use Cases

  • Image Generation: Create images from text descriptions using Stable Diffusion, SDXL, Flux, and other models.
  • Video Generation: Generate video content using AnimateDiff, Hunyuan, Wan, and other tools.
  • Audio Generation: Create audio content using various models and techniques.
  • Use Case: Imagine you want to create a short video of a fantasy landscape. Use this Skill to generate the images, compile them into a video, and add music to create a complete piece of content.

Quick Start

Generate an image with the prompt 'a fantasy landscape with mountains and a lake' using the ComfyUI skill.

Frequently Asked Questions about comfyui

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate images and video using ComfyUI?

You generate images and video using ComfyUI by providing text prompts to trigger models like Stable Diffusion or AnimateDiff. The Skill orchestrates this creative workflow, executing your requests through the REST and WebSocket API to produce the final media files.

Can I use Stable Diffusion and Flux models for image generation?

Yes, you can use Stable Diffusion, SDXL, and Flux models for image generation. The Skill leverages ComfyUI to apply these models to your text descriptions, creating custom visual content through the integrated API execution workflow.

What do I need to set up before generating audio content with ComfyUI?

To generate audio content with ComfyUI, you need the comfy-cli for environment setup and the associated models required for audio production. The Skill uses these dependencies to initialize the framework and execute your audio generation tasks via the API.

Does ComfyUI support video generation with AnimateDiff and Hunyuan?

Yes, ComfyUI supports video generation using AnimateDiff, Hunyuan, and Wan tools. The Skill manages this process by routing your text prompts to these specific models, compiling the generated frames into a complete video output.

What is the best way to create a short video of a fantasy landscape with audio?

The best way to create a short video of a fantasy landscape with audio is using ComfyUI for end-to-end creative workflows. You can generate the images, compile them into a video, and add music, all within a single automated execution pipeline.

Why does ComfyUI require the comfy-cli for setup?

ComfyUI requires the comfy-cli for setup to properly configure the environment and manage the associated models needed for media generation. This dependency ensures the REST and WebSocket API can successfully execute your image, video, and audio tasks.