What problem does it solve?
Converting reference images or video clips into high-quality, ready-to-use prompts for text-to-image and text-to-video generators is time-consuming and requires prompt engineering skill. This Skill automates visual analysis and crafts detailed prompts that capture subject, style, lighting, composition, and mood so creators can reproduce or remix visuals quickly.
Core Features & Use Cases
- Automated visual analysis: Extracts subject, style, color, lighting, composition, and atmosphere from images or videos.
- Multi-mode outputs: Produces prompts tailored for text-to-image, text-to-video, or both (image/video/auto modes).
- Language-aware prompts: Automatically chooses Chinese or English based on visual cues and text in the content.
- CLI and programmatic use: Includes a Python script with validation, streaming support, output saving, and error handling.
- Use Case: A designer uploads a reference photo and receives a polished Midjourney-ready prompt plus a breakdown explaining key prompt elements.
Quick Start
Ask the skill to analyze this image URL and produce a copy-paste-ready prompt for Midjourney that describes the subject, style, lighting, composition, color palette, and mood.