omnimedia

Analyze and generate audio, image, video, and document content via Gemini and MiniMax APIs.

5|Updated May 2, 2026
One-click install
npx skills add https://github.com/vanducng/skills --skill omnimedia
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: omnimedia
Source: https://github.com/vanducng/skills/tree/main/skills/omnimedia
Command: npx skills add https://github.com/vanducng/skills --skill omnimedia

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires google-genai, python-dotenv, Pillow, requests, pypdf, python-docx, docx2pdf, markdown, and includes scripts (resource) and references (resource) components.

What problem does it solve?

Automates multimodal analysis and generation across audio, images, videos, and documents using Gemini and MiniMax, enabling streamlined insight extraction and content production.

Core Features & Use Cases

  • Multimodal analysis: transcribe audio, OCR text, caption images, classify content, and extract structured data from diverse media formats.
  • Multimodal generation: create images, videos, speech, and music via Gemini, Imagen, OpenRouter, Codex (subscription), and MiniMax within unified workflows.
  • Document workflows: convert documents to Markdown, batch process files, and produce reusable artifacts for dashboards or knowledge bases.

Quick Start

Process a set of media files with Gemini to transcribe audio, describe images, and generate a short video.

Frequently Asked Questions about omnimedia

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I transcribe audio and extract text from images using multimodal AI?

Multimodal AI automates audio transcription and image OCR by processing diverse media formats through Gemini APIs to extract structured data and streamline insight extraction.

Can I generate images, videos, and speech within a unified workflow?

Yes, you can generate images, videos, speech, and music within unified workflows by leveraging Gemini, Imagen, OpenRouter, and MiniMax APIs for comprehensive content production.

What is the best way to convert documents to Markdown for knowledge bases?

Converting documents to Markdown is handled through document-to-Markdown workflows that batch process files, enabling reusable artifacts for dashboards or knowledge bases.

Does this multimodal generation approach work with PDF and DOCX files?

Yes, multimodal analysis processes PDF and DOCX files using dependencies like pypdf and python-docx to extract structured data and convert documents into Markdown format.

How do I batch process media files for transcription and video generation?

Batch processing of media files uses Gemini to transcribe audio, describe images, and generate short videos, enabling scalable media production across multiple file formats.

Do I need API keys for Gemini and MiniMax to run multimodal analysis?

Yes, multimodal analysis and generation require API keys for Gemini and MiniMax, managed through environment variables using python-dotenv for secure and streamlined access.