minimax-m3-multimodal-input

Ground multimodal inputs by reading media and citing exact file paths.

124|10|Updated Dec 8, 2025
One-click install
npx skills add https://github.com/madebyaris/advance-minimax-m3-cursor-rules --skill minimax-m3-multimodal-input
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: minimax-m3-multimodal-input
Source: https://github.com/madebyaris/advance-minimax-m3-cursor-rules/tree/main/.cursor/skills/minimax-m3-multimodal-input
Command: npx skills add https://github.com/madebyaris/advance-minimax-m3-cursor-rules --skill minimax-m3-multimodal-input

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Ground multimodal inputs (images and videos) to anchor visual claims in task reasoning.

Core Features & Use Cases

  • Ground visual claims by reading attached media and citing exact file paths.
  • Compare pre- and post-state media to verify UI or design changes.
  • Scope includes UI reviews, bug reports, and design parity checks with screenshots or clips.

Quick Start

Instruct the model to read the attached media and generate a grounded visual-fidelity verdict.

Frequently Asked Questions about minimax-m3-multimodal-input

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I ground visual claims from a screenshot during a UI review?

Ground visual claims by instructing the model to read attached media and cite exact file paths. The skill extracts on-screen text and produces a structured visual-fidelity verdict for your UI review.

What is the best way to verify design parity checks using pre and post state media?

Verify design parity by comparing pre- and post-state media. The skill anchors visual claims in reasoning tasks by reading the provided screenshots or short clips and generating a post-change comparison.

How do I extract exact file paths and on-screen text from a bug report video?

Extract exact file paths and on-screen text by applying multimodal grounding to the attached video. The skill reads the media content directly to anchor visual claims within your bug reports.

Does multimodal grounding work with both images and short video clips?

Yes, multimodal grounding works with both images and short video clips. The skill reads the attached media to ground visual claims and produces a structured visual-fidelity verdict for your design review.

Can I use this for design reviews without writing complex prompts?

Yes, you can use this for design reviews by simply instructing the model to read the attached media and generate a grounded visual-fidelity verdict. No complex prompt engineering is required.