image-vision

Analyze scenes, subjects, colors, text, and composition in user-attached images.

5|Updated Mar 14, 2026
One-click install
npx skills add https://github.com/chiptoe-svg/nanoclaw_gccourse --skill image-vision
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: image-vision
Source: https://github.com/chiptoe-svg/nanoclaw_gccourse/tree/main/container/skills/image-vision
Command: npx skills add https://github.com/chiptoe-svg/nanoclaw_gccourse --skill image-vision

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

When a user attaches an image to a conversation, they often need to understand the actual visual content of the photo — such as what is depicted, text present in the image, or the quality of the composition — rather than just basic file metadata like dimensions or file size. This Skill fills that gap by enabling AI assistants to actually "see" and analyze image content.

Core Features & Use Cases

  • Full Visual Content Analysis: Describe scenes, subjects, colors, mood, text, and composition of any attached image.
  • Custom Question Answering: Answer specific user questions about image content, such as critiquing an ad photo's effectiveness or checking if text on a sign is legible.
  • Adjustable Detail Levels: Choose between low (fast, low-cost, for overall scene overview), high (full resolution, for fine text or detail), or auto (model-selected) processing to balance speed, cost, and accuracy.
  • Use Case Example: A marketing professional attaches a draft social media ad photo and asks for feedback on its composition and visual appeal; this Skill analyzes the image and provides actionable critique.

Quick Start

Use the image-vision skill to analyze the attached product photo and tell me if the composition works for a social media ad.

Frequently Asked Questions about image-vision

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I analyze the visual content of an attached image?

To analyze visual content, attach the image to your conversation and request a description or critique. The AI "sees" the photo to identify scenes, subjects, colors, embedded text, and composition for full visual content analysis.

Can I extract text from images using AI?

Yes, you can extract text from images by attaching the file and asking for text extraction. Using the high detail level processes the image at full resolution, ensuring fine text and small details are accurately captured and read.

How do I critique a photo's composition for marketing assets?

Critique a photo's composition by attaching the marketing asset and requesting feedback. The AI assesses visual appeal, subjects, and layout, providing actionable critique on whether the composition works for social media ads.

Does image analysis support adjustable detail levels to balance cost and speed?

Yes, image analysis supports adjustable detail levels. Choose low for fast, low-cost scene overviews, high for full resolution fine detail extraction, or auto to let the vision model select the optimal balance.

What do I need to answer specific questions about image content?

To answer specific questions about image content, you need a vision-capable AI model and an attached image. You can then ask custom questions, such as checking if text on a sign is legible or evaluating visual effectiveness.

When should I use the high detail level for image description?

Use the high detail level for image description when you need full resolution analysis to read fine text or inspect intricate details. It balances cost and speed by processing the entire image at a higher resolution.