What problem does it solve?
This Skill simplifies the process of working with multimedia content, offering powerful AI-driven capabilities for audio, image, video, and document processing.
Core Features & Use Cases
- Audio Processing: Transcribe, analyze, and summarize audio files, with speech understanding and music/sound analysis.
- Image Understanding: Analyze images for captions, object detection, OCR, and visual Q&A.
- Video Analysis: Summarize, Q&A, and process video content with scene detection and temporal analysis.
- Document Extraction: Extract structured data from PDFs and convert documents to Markdown.
- Image Generation: Create images from text prompts, edit and modify existing images, and compose multiple images.
- Use Case: Imagine you need to process a large set of images, transcribe audio from a video, and extract key data from a PDF document. This Skill provides a comprehensive solution to handle all these tasks efficiently.
Quick Start
To analyze an image, use the following command: analyze_image "path/to/image.jpg" "describe this image" --model gemini-2.5-flash --output docs/assets/output.md