media-production

Automate AI visual media production with ZhipuAI APIs and FFmpeg.

12|3|Updated Jun 17, 2026
One-click install
npx skills add https://github.com/phuhao00/bony-agent --skill media-production-phuhao00
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: media-production
Source: https://github.com/phuhao00/bony-agent/tree/main/.agent/skills/media
Command: npx skills add https://github.com/phuhao00/bony-agent --skill media-production-phuhao00

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires requests, zhipuai, langchain, pillow, whisper, and includes scripts (resource) components.

What problem does it solve?

This Skill solves the complexity of managing fragmented media production workflows by unifying image generation, video synthesis, and professional-grade video editing into a single, automated pipeline.

Core Features & Use Cases

  • Multi-Modal Generation: Create high-quality images and videos from text prompts using advanced models like CogView-3 and CogVideoX.
  • Intelligent Remixing: Automatically analyze, sequence, and edit multiple media assets into a cohesive video production with transitions and motion effects.
  • Use Case: A content creator can provide a set of raw images and a creative theme, and the Skill will automatically generate a narrative script, synthesize motion-enhanced video segments, and assemble them into a final, ready-to-publish video.

Quick Start

Use the media-production skill to generate a 10-second promotional video based on the images in my uploads folder and the theme of a futuristic city.

Frequently Asked Questions about media-production

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate image generation and video assembly into a single workflow?

You can automate image generation and video assembly by using an end-to-end pipeline that synthesizes images from text prompts and assembles them with transitions. This skill unifies multi-asset sequencing and automated video assembly into a single workflow.

What is text-to-video conversion and how does intelligent remixing work?

Text-to-video conversion synthesizes video segments from text prompts, while intelligent remixing automatically analyzes, sequences, and edits multiple media assets into a cohesive video production with motion effects and transitions.

Do I need ZhipuAI and FFmpeg to run automated media production pipelines?

Yes, you need to integrate ZhipuAI APIs for advanced multi-modal generation and FFmpeg for high-performance media processing and transcoding. These dependencies are required to execute the automated media production pipeline.

Can I use existing images to generate a narrative script and motion-enhanced video?

Yes, you can provide a set of raw images and a creative theme. The skill will automatically generate a narrative script, synthesize motion-enhanced video segments using models like CogVideoX, and assemble them into a final video.

What are the limitations of using automated video remixing for complex creative workflows?

Automated video remixing relies on multi-asset sequencing and motion prompting, which may require precise text prompts and structured raw images. Complex creative workflows involving multi-asset sequencing depend heavily on ZhipuAI model capabilities and FFmpeg transcoding performance.