multimodal-looker

Analyze images and videos to extract descriptions, OCR text, and diagram insights.

12|4|Updated Jan 22, 2026
One-click install
npx skills add https://github.com/TurnaboutHero/oh-my-antigravity --skill multimodal-looker
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: multimodal-looker
Source: https://github.com/TurnaboutHero/oh-my-antigravity/tree/main/skills/multimodal-looker
Command: npx skills add https://github.com/TurnaboutHero/oh-my-antigravity --skill multimodal-looker

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps teams quickly understand visual content by providing structured analysis and descriptive summaries of images and videos, reducing manual inspection time.

Core Features & Use Cases

  • Image analysis and descriptive captions
  • UI/UX screenshot review and diagram interpretation
  • OCR text extraction from images and video frames
  • Video frame analysis for scene changes and key moments

Quick Start

Use the multimodal-looker skill to analyze a screenshot or video. Provide the file path or URL to the asset, and ask for a description or extracted text. For example: analyze image at /assets/ui-sample.png to get a UI overview.

Frequently Asked Questions about multimodal-looker

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract text from screenshots and images using OCR?

OCR text extraction analyzes images to pull out readable text content. This Skill performs OCR on screenshots, diagrams, and video frames, converting visual text into structured data without manual retyping.

Can I analyze UI screenshots to review design and layout?

Yes. This Skill provides detailed analysis of UI screenshots, identifying layout elements, visual hierarchy, and design patterns. It's built for design and QA workflows to assess interface clarity and usability.

What's the best way to interpret architecture diagrams and technical drawings?

Diagram interpretation reads and describes visual structures in technical drawings and architecture diagrams. This Skill extracts component relationships, flow paths, and system topology to help development teams understand complex visuals quickly.

How do I analyze video frames to identify key moments and scene changes?

Video frame analysis breaks video into individual frames and extracts descriptive insights about content, transitions, and significant moments. This Skill processes video to support design documentation, QA testing, and scene identification.

Does this work with both image files and video input?

Yes. This Skill handles both images and videos, extracting descriptions, text, and insights from either format. It supports UI screenshots, diagrams, and video frame analysis across your design and development workflows.