What problem does it solve?
This skill enables backend applications to understand and interact with images through natural language, removing the friction of building multimodal vision chat features and accelerating image analysis, OCR, classification, and visual Q&A workflows.
Core Features & Use Cases
- Vision chat integration: Combine images and text in multi-turn conversations to ask questions about images or request analyses.
- Flexible inputs: Accept image URLs and base64-encoded images, support multiple images, and handle image, video, and document attachments.
- Practical use cases: Product image descriptions for e-commerce, OCR and receipt extraction, accessibility alt text generation, image comparison for QA, and visual search.
Quick Start
Ask the system to describe the image at https://example.com/photo.jpg and list the main objects and any text present.