multimodal-looker

Extract text, entities, and relationships from images, PDFs, and charts.

16|1|Updated Jan 2, 2026
One-click install
npx skills add https://github.com/bahayonghang/my-claude-code-settings --skill multimodal-looker-bahayonghang
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: multimodal-looker
Source: https://github.com/bahayonghang/my-claude-code-settings/tree/main/skills/multimodal-looker
Command: npx skills add https://github.com/bahayonghang/my-claude-code-settings --skill multimodal-looker-bahayonghang

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

The Multimodal Looker helps users interpret and extract structured information from visual content (images, PDFs, charts, diagrams) without manual sifting, saving time and reducing ambiguity.

Core Features & Use Cases

  • Visual content analysis: extract text, identify UI elements, and describe layouts from screenshots.
  • PDF data extraction: pull text, tables, and metadata from documents.
  • Chart and diagram interpretation: deduce data series, relationships, and data flow from charts, diagrams, and design drafts.
  • Context-efficient reporting: convert visuals into concise, structured text summaries for further processing.

Quick Start

Provide a visual input (image or PDF) and a task such as "summarize this chart" or "extract all text and key entities," then run multimodal-looker to receive a structured description of components, relationships, and extracted data.

Frequently Asked Questions about multimodal-looker

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract data from a PDF chart or image for structured analysis?

You can extract data from a PDF chart or image by providing the visual input and a task instruction. The tool then outputs a structured description of components, relationships, and extracted text.

Can I extract text and identify key entities from screenshots?

Yes, you can extract text and identify key entities from screenshots. The analysis identifies UI elements and describes layouts to produce structured, actionable insights without manual sifting.

What is the best way to interpret diagrams and generate a data summary?

The best way to interpret diagrams and generate a data summary is to run the analysis on your visual input. It deduces data series and relationships to output concise text summaries.

Does this visual content analysis approach work for technical audits and design reviews?

Yes, visual content analysis works for technical audits and design reviews. It interprets diagrams and design drafts to provide extracted text, key entities, and data flow relationships.

How do I convert visual insights into concise natural language outputs?

You convert visual insights into natural language outputs by processing images or PDFs through multimodal analysis, which generates context-aware explanations and structured summaries.

Are there limitations when extracting tables and metadata from PDF documents?

Limitations when extracting tables and metadata from PDF documents include relying on clear visual inputs, as the tool deduces data relationships without exposing or processing sensitive underlying data.