vocabulary-video-pipeline

Automate vocabulary video production from word to Feishu upload.

225|37|Updated Apr 3, 2026
One-click install
npx skills add https://github.com/dracohu2025-cloud/draco-skills-collection --skill vocabulary-video-pipeline
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: vocabulary-video-pipeline
Source: https://github.com/dracohu2025-cloud/draco-skills-collection/tree/main/vocabulary-video-pipeline
Command: npx skills add https://github.com/dracohu2025-cloud/draco-skills-collection --skill vocabulary-video-pipeline

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires pydub, and includes scripts (resource) and assets (resource) components.

What problem does it solve?

This Skill automates the end-to-end process of turning a single English word into a polished vocabulary video, combining diagnosis, TTS, beat-based timing, Remotion rendering, uploading, and cost reporting.

Core Features & Use Cases

  • End-to-end pipeline: diagnose word data, synthesize audio, segment beats, render with Remotion, and upload the result.
  • Template-driven storytelling: supports origin-chain and multiple scene templates to produce engaging educational content.
  • Reproducible workflow: generates draft outputs and cost logs for transparent budgeting and iteration.

Quick Start

Provide a word and start the automated workflow to generate a draft, audio, and final video in one run.

Frequently Asked Questions about vocabulary-video-pipeline

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate vocabulary video production from a single word?

Automating vocabulary video production requires providing a single word to trigger an end-to-end pipeline that diagnoses data, synthesizes TTS audio, segments beats, renders video, and uploads the result automatically.

How does beat-aware timing work when generating educational videos?

Beat-aware timing works by validating beats during the workflow to sync Remotion visual rendering with Volcengine TTS audio, ensuring visual scene transitions match the synthesized speech rhythm for engaging educational content.

Can I use Remotion with Volcengine TTS to generate educational content?

Yes, you can use Remotion with Volcengine TTS to generate educational content by synthesizing speech, segmenting beats, and rendering template-driven videos through an automated pipeline that enforces required validation steps.

Do I need pydub to render video and upload to Feishu?

You need pydub as a dependency to process audio within the pipeline, which subsequently renders the video using Remotion and handles the automatic upload of the finished vocabulary video to Feishu.

What is the origin-chain template in vocabulary video workflows?

The origin-chain template is a required storytelling format within the workflow that structures the educational content, ensuring the vocabulary video follows a validated narrative scene progression before director signoff.

How do I track rendering costs for TTS and Remotion video generation?

Tracking rendering costs for TTS and Remotion video generation is handled automatically by the pipeline, which generates transparent cost logs and draft outputs to provide reproducible budgeting for your educational content.