llava

Describe images and answer visual questions in conversational dialogue.

Updated Jun 19, 2026
One-click install
npx skills add https://github.com/AnandaAnugrahHandyanto/savarez_agent --skill llava-anandaanugrahhandyanto
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: llava
Source: https://github.com/AnandaAnugrahHandyanto/savarez_agent/tree/main/optional-skills/mlops/llava
Command: npx skills add https://github.com/AnandaAnugrahHandyanto/savarez_agent --skill llava-anandaanugrahhandyanto

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Provides a multimodal AI assistant capable of understanding and describing images and engaging in visual-based conversations.

Core Features & Use Cases

  • Multimodal reasoning that combines a vision encoder with a language model for image-centric dialogue
  • Visual question answering and image captioning in multi-turn conversations
  • Visual instruction-following and image understanding across domains such as accessibility, customer support, and data analysis
  • Easy integration with vision-language tasks and custom datasets

Quick Start

Provide an image and request a multimodal analysis, including description and questions.

Frequently Asked Questions about llava

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I do visual question answering with an image in a conversational AI?

Visual question answering combines a vision encoder with a language model to process images and respond to queries. This Skill handles multimodal dialogue, allowing you to provide an image and request descriptions or specific answers about its content.

Can I use this for image captioning and multi-turn visual dialogue?

Yes, image captioning and multi-turn visual dialogue are core features. The multimodal reasoning engine understands visual instructions and maintains context across conversational turns to describe images and answer follow-up questions.

What is multimodal visual instruction following for image-based chat?

Multimodal visual instruction following is a process where a vision-language model interprets images and executes text-based commands. It enables AI to analyze visual content and engage in image-centric chat across domains like accessibility and data analysis.

How do I integrate vision-language tasks with custom datasets?

You can integrate vision-language tasks by providing custom images and specific visual instructions within the chat. The Skill processes the visual inputs against your custom datasets to generate tailored multimodal analysis and conversational responses.

Does visual instruction following work for accessibility and customer support use cases?

Visual instruction following works well for accessibility and customer support by analyzing image content and responding to contextual queries. It applies multimodal reasoning to understand visual data and provide descriptive answers in a conversational format.