What problem does it solve?
It turns raw text and media into the ready-to-composite assets HyperFrames compositions need, saving you from manual narration, caption timing, and background cutout workflows.
Core Features & Use Cases
- Text-to-speech narration (TTS): Generate spoken audio locally with Kokoro, choose from many voices, and cache models on first run.
- Audio/video transcription with word timestamps: Produce a normalized transcript.json for caption overlays and downstream use.
- Background removal for transparent overlays: Create cutout foreground (and an optional inverse “plate” layer) for transparent video/photo compositing.
- Use Case: Generate a voiceover from a script, transcribe it to get word-level timestamps for captions, and remove the background from a presenter clip to place them over custom graphics.
Quick Start
Use hyperframes-media to create narration, transcribe it, and then generate transparent cutout assets by running it as a local-first pipeline.