What problem does it solve?
LLaVA removes the gap between language models and visual input by letting you ask questions about images, generate descriptions, and hold multi-turn conversations grounded in what is seen.
Core Features & Use Cases
- Visual Question Answering: Answer detailed questions about photos, screenshots, and documents.
- Image Captioning and Scene Understanding: Produce concise captions or rich scene summaries for single images.
- Multi-Turn Multimodal Chat: Keep context across follow-up questions about the same image.
- Training and Fine-Tuning Support: Learn how to pretrain, instruction-tune, quantize, and adapt LLaVA for custom datasets.
- Use Case: Turn a product screenshot into a troubleshooting assistant that explains what is visible and answers iterative questions from a support agent.
Quick Start
Use the llava skill to analyze the attached image, describe what it shows, and answer any follow-up visual questions in plain language.