ck:ai-multimodal

Analyze and generate images, audio, video, speech, and music using AI models.

1|Updated May 4, 2026
One-click install
npx skills add https://github.com/auxi-wardrobe/auxi-all-in --skill ck-ai-multimodal-auxi-wardrobe
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ck:ai-multimodal
Source: https://github.com/auxi-wardrobe/auxi-all-in/tree/main/.agents/skills/ai-multimodal
Command: npx skills add https://github.com/auxi-wardrobe/auxi-all-in --skill ck-ai-multimodal-auxi-wardrobe

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires google-genai, requests, PIL, requests, and includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill streamlines the process of analyzing and generating multimedia content by leveraging powerful AI models, enabling users to efficiently process images, audio, and video.

Core Features & Use Cases

  • Multimedia Analysis: Analyze images, audio, and video using Gemini API for advanced vision, audio, and video analysis.
  • Image Generation: Generate images via Google, OpenRouter, or MiniMax models.
  • Video Generation: Generate videos using Veo models.
  • Speech Generation: Generate speech (TTS) using MiniMax models.
  • Music Generation: Generate music using MiniMax models.
  • Use Case: A designer can use this Skill to generate concept art, analyze fashion trends from images, or create marketing videos.

Quick Start

Analyze an image using the ai-multimodal skill with the command: analyze_image <image_path>.

Frequently Asked Questions about ck:ai-multimodal

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate images and videos using the Gemini API?

This Skill automates multimedia generation using the Gemini API, Veo models, and MiniMax, enabling you to generate images and videos directly from text prompts for design workflows.

Can I analyze images, audio, and video content using AI models?

Yes, the Skill performs advanced AI analysis on images, audio, and video using the Gemini API, extracting visual and auditory insights for multimedia content processing workflows.

Do I need Google GenAI, OpenRouter, and MiniMax API keys to generate speech and music?

Yes, you must provide Google GenAI, OpenRouter, and MiniMax API keys to authenticate and access the underlying models for speech generation, music generation, and multimedia creation.

What is the best way to automate multimedia analysis and generation for design workflows?

The best way to automate multimedia analysis and generation is using this Skill, which integrates Gemini API for content analysis and Veo or MiniMax models for content creation.

How do I analyze an image to extract fashion trends using AI?

You can analyze an image to extract fashion trends by passing the file path to the Skill's image analysis function, which uses the Gemini API for advanced vision analysis.