vision

Extract text, classify objects, and analyze charts from images.

Updated Nov 7, 2025
One-click install
npx skills add https://github.com/Wesley1600/ClaudeCodeFrameWork --skill vision-wesley1600
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: vision
Source: https://github.com/Wesley1600/ClaudeCodeFrameWork/tree/main/.claude/skills/vision
Command: npx skills add https://github.com/Wesley1600/ClaudeCodeFrameWork --skill vision-wesley1600

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires Pillow, and includes scripts (resource) and references (resource) components.

What problem does it solve?

Analyzes and processes images to extract text, classify content, analyze charts, compare diagrams, and describe visuals, enabling automated understanding of visual data.

Core Features & Use Cases

  • OCR/text extraction from images and receipts
  • Image classification and object detection
  • Chart/diagram analysis and visual QA
  • Image description and accessibility-friendly outputs

Quick Start

Provide an image file to analyze: e.g., analyze.png to extract text and insights.

Frequently Asked Questions about vision

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract text from images using OCR?

OCR extracts text from images by analyzing pixel patterns to recognize characters. This Skill applies Claude's vision capabilities to extract text from screenshots, photos, receipts, and scans, returning results in Markdown format with confidence annotations.

Can I analyze charts and diagrams to extract data?

Yes. This Skill interprets visual data in charts, diagrams, and graphs by identifying structure, labels, and values. It supports chart analysis and diagram comparison, preserving layout information in the output.

What image formats and sources does this support?

The Skill processes images including screenshots, photos, and scans through multi-modal input handling. It supports common image formats and applies visual analysis across all input types while maintaining structure in outputs.

How do I classify objects or identify content in images?

Image classification identifies and categorizes visual content by analyzing objects, patterns, and context. This Skill detects objects and classifies image content, returning descriptions with supported confidence metrics.

Can I use this for document processing and accessibility?

Yes. This Skill processes document images, extracts structured text, and generates accessibility-friendly descriptions. It supports multiple languages and formats outputs for readability and assistive technology compatibility.

What does visual Q&A mean and how does it work?

Visual Q&A answers questions about image content by analyzing visual elements and context. This Skill interprets images to answer specific queries about what they contain, supporting UI/UX assessment and detailed visual analysis.