What problem does it solve?
This Skill helps you turn natural-language ideas, reference images, and compositing needs into high-quality Grok image or video outputs without guessing the right API shape.
Core Features & Use Cases
- Text-to-Image: Create standalone visuals from prompts with model, aspect ratio, and resolution controls.
- Image Editing and Compositing: Modify one image or combine multiple sources while preserving the intended subject and structure.
- Text/Image-to-Video: Produce short clips from prompts, an anchor image, or multiple reference images for product demos, fashion transfer, and character-consistent motion.
- Operational Guidance: Handle authentication, request submission, and async polling so generation jobs can be monitored to completion.
Quick Start
Use this skill to generate a square Grok image from the prompt "a cinematic product shot of a silver watch on black stone".