What problem does it solve?
This Skill helps agents extract reliable, structured information from still images without relying on open-ended visual reasoning. It supports production workflows that need searchable labels, object locations, OCR, moderation signals, color data, crop suggestions, and web-reference evidence.
Core Features & Use Cases
- Structured Image Analysis: Detect labels, objects, dominant colors, crop hints, explicit-content likelihoods, and web matches.
- OCR and Document Processing: Choose sparse-text detection for signs, labels, and screenshots or document-text detection for dense pages, handwriting, PDFs, and TIFFs.
- Production Pipelines: Design synchronous and asynchronous Cloud Storage workflows with IAM, quotas, cost estimation, regional OCR endpoints, retention controls, retries, and quality evaluation.
- Use Case: Process a large media library with searchable tags, object boxes, poster-text OCR, and SafeSearch-based review routing while preserving auditability and human oversight.
Quick Start
Use the google-cloud-vision skill to analyze the uploaded still images for labels, object locations, poster text, explicit-content review signals, and production QA recommendations.