multimodal-looker

Analyze images, screenshots, diagrams, and document visuals into structured observations and recommendations.

10|Updated Mar 22, 2026
One-click install
npx skills add https://github.com/Lee-SiHyeon/oh-my-copilot --skill multimodal-looker-lee-sihyeon
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: multimodal-looker
Source: https://github.com/Lee-SiHyeon/oh-my-copilot/tree/main/.github/skills/multimodal-looker
Command: npx skills add https://github.com/Lee-SiHyeon/oh-my-copilot --skill multimodal-looker-lee-sihyeon

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

It helps you understand visual content (screenshots, diagrams, and PDFs) when the meaning, structure, or extracted details are not obvious from a quick glance.

Core Features & Use Cases

  • Visual analysis for UI and layouts: Break down layout, components, hierarchy, typography, and concrete visual issues in screenshots or mockups.
  • Diagram and architecture interpretation: Identify entities, relationships, and data/control flow to explain how the system works.
  • Document summarization and extraction: Summarize key points and extract structured details like tables, numbers, and action items from documents.

Quick Start

Ask the assistant: "Analyze this screenshot and extract the key UI elements, visual issues, and specific recommendations to improve it."

Frequently Asked Questions about multimodal-looker

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract key UI elements and visual issues from a screenshot?

Screenshot analysis breaks down layout, components, hierarchy, and typography to identify visual issues. It produces structured output covering observations, analysis, key findings, and actionable recommendations to improve the interface.

Can I interpret system architecture and data flow from a diagram image?

Diagram interpretation identifies entities, relationships, and data/control flow to explain how the system works. It analyzes visual context to produce structured explanations covering observations, analysis, key findings, and recommendations.

What is the best way to summarize a PDF and extract structured data like tables?

Document summarization analyzes document visuals to extract structured details like tables, numbers, and action items. It interprets visual context to summarize key points and produce structured output covering observations, analysis, and key findings.

Does multimodal image analysis work for read-only document and screenshot review?

Yes, multimodal image analysis supports read-only analysis for screenshots and documents. It interprets visual context without modifying the original file, producing structured explanations and actionable findings from the visual data.

What format does the visual analysis output use for findings and recommendations?

Visual analysis output uses a structured format covering observations, analysis, key findings, and recommendations. This ensures actionable insights extracted from images, diagrams, and documents are clearly organized and accessible.