What problem does it solve?
This Skill eliminates the manual work of producing all audio and visual media assets required for HyperFrames video compositions, including voiceover, background music, sound effects, transcription, timed captions, and subject background removal, which would otherwise require integrating multiple separate tools and services.
Core Features & Use Cases
- Multi-provider Text-to-Speech: Supports HeyGen (with native word timestamps), ElevenLabs, and offline local Kokoro TTS, with automatic fallback based on available credentials and language needs.
- Background Music & Sound Effects: Retrieves royalty-free BGM and SFX from HeyGen's library when credentialed, or falls back to local Lyria/MusicGen BGM generation and a bundled offline SFX library when no credentials are available.
- Transcription & Caption Authoring: Generates word-level timestamps for TTS audio, cleans raw transcripts, and creates customizable captions with karaoke effects, motion animations, and audio-reactive modulation.
- Background Removal: Creates transparent cutouts of subjects from video or image files for compositing into custom scenes.
- Use Case: For a product launch video, use this Skill to generate a polished voiceover from a script, add matching background music, insert UI interaction sound effects, auto-generate perfectly timed captions, and cut out the presenter to overlay on a branded background.
Quick Start
Use the hyperframes-media skill to produce all required audio, caption, and background removal assets for your HyperFrames video composition from a provided script and scene cues.