What problem does it solve?
This Skill removes the friction of manually inspecting media files by turning images, audio, video, and documents into searchable insights, transcripts, and generated assets.
Core Features & Use Cases
- Multimodal Analysis: Inspect images for OCR, classification, object detection, scene understanding, and document extraction.
- Audio and Video Workflows: Transcribe speech, summarize videos, detect scenes, and analyze long recordings with timestamps and speaker context.
- Generation Workflows: Create images, videos, speech, and music with Gemini and MiniMax models for creative and production use cases.
- Document Conversion: Convert PDFs, Office files, and web content into clean Markdown for knowledge capture and reuse.
- Use Case: A product team can upload screenshots, meeting audio, and demo videos, then extract text, summarize feedback, and generate polished marketing visuals from the same workflow.
Quick Start
Ask the skill to analyze the attached media file or generate a new image, video, speech, music, or document summary using the appropriate Gemini or MiniMax model.