multimodal-llm

Process images, transcribe audio, generate speech, and create videos with multimodal AI models.

217|20|Updated Dec 31, 2025
One-click install
npx skills add https://github.com/yonatangross/orchestkit --skill multimodal-llm
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: multimodal-llm
Source: https://github.com/yonatangross/orchestkit/tree/main/plugins/ork/skills/multimodal-llm
Command: npx skills add https://github.com/yonatangross/orchestkit --skill multimodal-llm

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill enables seamless integration of advanced multimodal AI capabilities, allowing you to process images, transcribe audio, generate speech, and create AI-generated video content.

Core Features & Use Cases

  • Image Analysis: Understand and describe images, extract data from documents and charts.
  • Audio Processing: Transcribe speech to text, generate natural-sounding speech from text.
  • Video Generation: Create AI-powered videos using cutting-edge models like Kling, Sora, and Veo.
  • Use Case: Build an AI assistant that can describe images uploaded by users, transcribe meeting recordings, and generate short promotional videos for products.

Quick Start

Use the multimodal-llm skill to describe the provided image and generate a short video based on a text prompt.

Frequently Asked Questions about multimodal-llm

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate AI video from text while keeping character consistency across multiple shots?

You can perform AI video generation with character consistency across multiple shots by using this multimodal integration, which leverages models like Kling, Sora, and Veo to maintain visual continuity throughout the sequence.

Can I transcribe speech to text and generate natural sounding speech from text in one workflow?

Yes, you can transcribe speech to text and generate natural-sounding speech from text within a single workflow using the audio processing capabilities of this multimodal AI integration.

What is the best way to extract data from documents and charts using image analysis?

The best way to extract data from documents and charts is through multimodal image analysis, which performs detailed document OCR to accurately pull structured information from visual inputs.

Does this multimodal AI approach support both image understanding and video generation tasks?

Yes, this multimodal AI approach supports both image understanding and video generation tasks, allowing you to analyze uploaded images and create AI-generated videos using models like Kling, Sora, and Veo.

How do I build an AI assistant that can process images, transcribe meetings, and generate videos?

You can build an AI assistant that processes images, transcribes meetings, and generates videos by integrating this multimodal AI skill, which unifies image analysis, speech-to-text transcription, and AI video generation.