image-analyzer

Extract text, describe scenes, and detect objects from images.

12|4|Updated Feb 12, 2026
One-click install
npx skills add https://github.com/orthogonal-sh/skills --skill image-analyzer
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: image-analyzer
Source: https://github.com/orthogonal-sh/skills/tree/main/skills/orthogonal-image-analyzer
Command: npx skills add https://github.com/orthogonal-sh/skills --skill image-analyzer

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Analyze images to extract text, describe content, and detect objects with AI.

Core Features & Use Cases

  • OCR: Extract text from images for accessibility and data capture.
  • Image description: Generate natural-language descriptions of visual scenes.
  • Object detection: Identify and locate objects within images for cataloging.

Quick Start

Provide an image URL or upload an image and receive text extraction, description, and object detection results.

Frequently Asked Questions about image-analyzer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract text from an image using OCR?

Extracting text from an image using OCR requires providing an image URL or upload to the analyzer. The tool processes the visual data and returns structured text outputs suitable for downstream data collection and accessibility workflows.

What is AI-driven image description and object detection?

AI-driven image description and object detection is the process of generating natural-language scene summaries and identifying specific items within an image. It outputs structured visual insights applicable to cataloging, marketing, and accessibility workflows.

Can I process batches of image URLs for content extraction?

Yes, you can process batches of image URLs for content extraction. The analysis supports batches to perform OCR, scene description, and object detection across multiple images simultaneously for data-collection workflows.

Does image analysis require a specific backend for computer vision capabilities?

Yes, performing image analysis requires an image processing backend with OCR and computer vision capabilities. These backend components are necessary to detect objects, describe scenes, and extract text from provided images.

What is the best way to structure visual insights for downstream AI prompts?

The best way to structure visual insights for downstream AI prompts is using an image analyzer that formats extracted text, scene descriptions, and detected objects into structured outputs. This ensures compatibility with subsequent data processing.

When do I need image analysis for accessibility workflows?

You need image analysis for accessibility workflows when you must extract visible text from images or generate natural-language descriptions of visual scenes. This enables screen readers and assistive technologies to convey visual content to users.