What problem does it solve?
Producing AI-generated images, videos, and digital-human content on Aliyun Bailian requires juggling many model endpoints, protocols, and billing rules; this Skill unifies them into one studio with 22 LLM-callable tools and a storyboard-driven long-video pipeline.
Core Features & Use Cases
- Unified image and video generation: Text-to-image, image edit, style repaint, background generation, outpainting, sketch-to-image, and ecommerce scenes, plus text-to-video, image-to-video, reference-to-video, and video editing via HappyHorse 1.0 and Wan 2.6/2.7 models.
- Digital-human modes: Photo speak, video relip, video reface, pose drive, and avatar compose, with CosyVoice-v2 and Edge-TTS for voice synthesis.
- Storyboard long-video pipeline: Decompose a story into structured segments, generate them serially or in parallel, and concat with ffmpeg crossfade transitions.
- Use Case: Ask the agent to turn a short story into a 30-second video; it decomposes the story into storyboard segments, generates each clip, and concatenates them into a final MP4 with cost preview before submission.
Quick Start
Ask the agent to generate a 5-second 720P text-to-video clip of a sunrise over the sea using the happyhorse-video tools after configuring your DashScope API key and OSS settings.