What problem does it solve?
Aligning footage from multiple cameras and separate audio recorders covering one shoot is tedious and error-prone when done by hand. This Skill analyzes transcripts across all clips, computes precise timeline offsets, identifies who is on each camera and microphone, and optionally cuts a speaker-follow draft — all without creating fragile multicam clip objects.
Core Features & Use Cases
- Automatic structure discovery: Reads every clip's transcript to determine how many cameras and recorders exist, distinguishing simultaneous angles from sequential spans of one camera.
- Deterministic offset calculation: Uses a bundled script that votes on shared phrases and takes the median of matched utterance pairs to compute sub-frame-accurate sync offsets, with confidence flags and drift detection.
- Evidence-based labeling: Identifies which camera frames whom and which mic belongs to which speaker using frame inspection and in-track loudness contrast, then labels tracks accordingly.
- Speaker-follow cutting: On request, builds a watchable draft on a separate timeline that cuts to the active speaker while keeping program audio continuous and the synced master untouched.
- Use Case: You filmed a two-person interview with two cameras and two lav recorders. Drop in all the clips, and the Skill aligns them, labels each track by speaker, and produces a draft edit that follows the conversation.
Quick Start
Sync these camera and audio clips from my interview shoot and cut a draft that follows whoever is speaking.