llava

Answer questions about images using CLIP-based vision encoding and Vicuna/LLaMA language models.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/Hermesagents/hermes-agents --skill llava-hermesagents
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: llava
Source: https://github.com/Hermesagents/hermes-agents/tree/main/skills/mlops/models/llava
Command: npx skills add https://github.com/Hermesagents/hermes-agents --skill llava-hermesagents

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Provides a unified vision-language capability that lets users ask questions about images and receive coherent, context-aware responses.

Core Features & Use Cases

  • Multimodal image conversation: discuss visuals across multiple turns.
  • Visual question answering (VQA): answer questions about objects, scenes, and actions in images.
  • Image captioning and description: generate natural language descriptions of visual content.
  • Instruction following with visual context: perform tasks based on images and prompts.

Quick Start

Upload an image and start a multimodal chat to get descriptive, analytical, or instruction-driven responses.

Frequently Asked Questions about llava

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How does visual question answering work for image chat?

Visual question answering for image chat works by applying CLIP-based vision encoding with Vicuna or LLaMA-style language models to answer questions about objects, scenes, and actions in images. It supports multi-turn dialogs with visual context.

What is multimodal instruction-following for visual content?

Multimodal instruction-following for visual content is a vision-language task where you upload an image and provide prompts to receive coherent, context-aware responses. It lets you perform tasks based on images and instructions across iterative conversations.

Can I use Vicuna and LLaMA models for multi-turn image conversations?

Yes, you can use Vicuna and LLaMA-style language models for multi-turn image conversations. The system relies on CLIP-based vision encoding combined with these language models to support iterative conversations with visual context.

How do I start an image chat to generate descriptions of visual content?

To start an image chat and generate descriptions of visual content, simply upload an image and begin a multimodal chat. You will receive descriptive, analytical, or instruction-driven responses about the visual content.

What is the best way to perform visual instruction-following across multiple turns?

The best way to perform visual instruction-following across multiple turns is to use a unified vision-language capability that maintains visual context. This approach lets you ask questions and give prompts about images over iterative conversations.

Does this image chat approach support CLIP-based vision encoding with language models?

Yes, this image chat approach supports CLIP-based vision encoding integrated with Vicuna or LLaMA-style language models. This combination enables context-aware responses for visual question answering and image captioning tasks.