What problem does it solve?
This Skill provides a unified interface to analyze and generate multimedia content across audio, image, video, and document formats using Google's Gemini multimodal APIs. It enables automated transcription, captioning, object detection, OCR, video scene analysis, and image generation, reducing manual effort and enabling multimodal AI workflows.
Core Features & Use Cases
- Audio processing: transcription with timestamps, summarization, speech understanding, and TTS where applicable.
- Image understanding: captioning, object detection with bounding boxes, segmentation, OCR, and visual Q&A.
- Video analysis: scene detection, Q&A, temporal analysis, YouTube URL support, long-form video processing.
- Document extraction: tables, forms, charts from PDFs; multi-page understanding; data extraction.
- Image generation: text-to-image, editing, composition, refinement with various aspect ratios and styles.
- Works with multiple Gemini models (Gemini 2.5/2.0) with large context windows.
Quick Start
Use the ai-multimodal skill to transcribe an audio file, summarize it, caption images, and generate a related image from a prompt.