ag-capturar-tela

Captures screen regions across macOS, Windows, and Linux for multimodal AI analysis and OCR extraction.

19|4|Updated Mar 7, 2026
One-click install
npx skills add https://github.com/andregusman-raiz/a-gusman-claude --skill ag-capturar-tela
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ag-capturar-tela
Source: https://github.com/andregusman-raiz/a-gusman-claude/tree/main/skills/ag-capturar-tela
Command: npx skills add https://github.com/andregusman-raiz/a-gusman-claude --skill ag-capturar-tela

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This utility captures the current screen across macOS, Windows, and Linux and analyzes visual content using multimodal AI to describe what is visible, read text via OCR, and identify the active apps.

Core Features & Use Cases

  • Cross-platform screen capture for native apps, terminals, and multiple windows.
  • Multimodal analysis including visual description and OCR to extract visible text.
  • Use case: quickly understand what a user sees on their screen or extract important text from a terminal output.

Quick Start

Invoke ag-capturar-tela to capture and describe your current screen, or specify a region or monitor to focus the analysis.

Frequently Asked Questions about ag-capturar-tela

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I extract text from a screen capture using AI?

You can extract text from a screen capture using built-in multimodal AI and OCR capabilities. The tool analyzes visible screen elements across desktop environments to read and extract text from native apps or terminal output.

Does cross-platform screen capture work on macOS, Windows, and Linux?

Cross-platform screen capture works natively on macOS, Windows, and Linux. It identifies active windows and native apps across these desktop environments to ensure consistent visual analysis.

Can I target a specific monitor or screen region for visual analysis?

Yes, you can target a specific monitor or screen region for visual analysis. Instead of capturing the entire desktop, you specify the focus area to isolate visible elements and extract relevant text.

What is multimodal screen analysis and how does it describe on-screen content?

Multimodal screen analysis uses AI to describe on-screen content by interpreting visual elements. It identifies active windows, native apps, and terminal text, providing a comprehensive description of what is currently visible.

How do I capture and analyze terminal output from my desktop?

To capture and analyze terminal output, invoke the screen capture utility in your desktop environment. It uses OCR and visual description to quickly extract and understand important text from the terminal window.

Are there limitations when identifying active windows in desktop environments?

Limitations when identifying active windows depend on the underlying desktop environment permissions. Screen capture requires proper accessibility access on macOS, Windows, or Linux to accurately identify native apps and read visible text.