ck:ai-multimodal

Analyze and generate images, videos, audio, and text using Gemini and MiniMax APIs.

1|Updated Jun 16, 2026
One-click install
npx skills add https://github.com/TNHoang2708/Gym_Ver2 --skill ck-ai-multimodal-tnhoang2708
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ck:ai-multimodal
Source: https://github.com/TNHoang2708/Gym_Ver2/tree/main/.claude/skills/ai-multimodal
Command: npx skills add https://github.com/TNHoang2708/Gym_Ver2 --skill ck-ai-multimodal-tnhoang2708

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires google-genai, python-dotenv, pillow, requests, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill provides a comprehensive suite of AI tools for analyzing and generating images, videos, audio, and text, streamlining tasks like image analysis, video generation, and text transcription.

Core Features & Use Cases

  • Image Analysis: Analyze images for content, extract text, and perform OCR.
  • Video Generation: Create videos from text descriptions or images.
  • Audio Processing: Transcribe audio, generate speech, and analyze audio content.
  • Text Generation: Generate images, videos, and audio from text descriptions.
  • Use Case: Imagine you need to create a promotional video for a new product. Use this Skill to generate a video from a text description, analyze the product images for key features, and transcribe the audio for accessibility.

Quick Start

Use the ai-multimodal skill to generate a video from the text description 'A promotional video showcasing the latest smartphone model in an urban environment with a focus on its camera capabilities.'

Frequently Asked Questions about ck:ai-multimodal

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate a video from a text description using AI?

To generate a video from text descriptions, you can use this Skill to process your prompts via the MiniMax API, creating visual content directly from written input.

Can I perform OCR and extract text from images with AI tools?

Yes, you can perform OCR and extract text from images using this Skill's image analysis capabilities, which leverage the Gemini API to analyze and process image content.

Does this multimodal AI tool work with the Gemini API and MiniMax API?

Yes, this Skill integrates directly with both the Gemini API and MiniMax API to provide a wide range of AI capabilities for image, video, audio, and text processing.

How do I transcribe audio and generate speech from text?

You can transcribe audio and generate speech from text by utilizing this Skill's audio processing features, which handle both audio content analysis and speech generation.

What Python dependencies are needed to run AI multimodal processing?

You need the google-genai, python-dotenv, pillow, and requests Python dependencies installed in your environment to run this AI multimodal processing Skill.