captions-overlay

Applies caption overlay doctrine for compositing subtitles on video compositions.

43.5k|4.2k|Updated Mar 10, 2026
One-click install
npx skills add https://github.com/heygen-com/hyperframes --skill captions-overlay
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: captions-overlay
Source: https://github.com/heygen-com/hyperframes/tree/main/.agents/skills/captions-overlay
Command: npx skills add https://github.com/heygen-com/hyperframes --skill captions-overlay

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

When adding captions to talking-head or launch videos, creators often misplace content by reserving a dead bottom band for subtitles or overusing embedded word effects, producing unbalanced compositions. This Skill defines the caption model and overlay rules so captions composite cleanly on top of the film without distorting the layout.

Core Features & Use Cases

  • Caption Model (drop / rail / embed): Classifies every spoken phrase as dropped filler, verbatim lower-third rail text, or a scarce embedded peak word composited behind the subject.
  • Overlay Law Enforcement: Ensures captions are treated as an overlay layer on the true vertical center of the frame, never a reserved keep-out band that shifts content upward.
  • Use Case: When building a product launch video with word-by-word captions, use this Skill to center the composition at y = H/2, keep the verbatim rail readable, and promote at most one word per beat to an embedded climax.

Quick Start

Use the captions-overlay skill to add rail-first captions with one embedded peak word to my talking-head video composition.

Frequently Asked Questions about captions-overlay

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I add captions to a talking-head video without breaking the layout?

Treat captions as an overlay composited on top of the film, not a reserved bottom band. Center the composition on the true vertical center (y = H/2) and let the verbatim rail carry most spoken text as a lower-third subtitle.

What is the drop, rail, and embed caption model?

Every spoken phrase is classified as drop (filler, not shown), rail (verbatim lower-third subtitle, the default), or embed (a scarce peak word composited behind the subject). The rail carries most text; embeds are limited to one per beat.

Should I reserve a bottom band in my video layout for captions?

No. Captions are an overlay layer, not a reserved zone, so content may extend to the canvas bottom. Only avoid parking critical small readable text like URLs in the bottom ~80px center span where the caption line sits.

When should I use embedded caption words instead of the rail?

Use embed only for genuine peak moments, at most one per sentence or beat, never two adjacent or co-visible. Embedding every word is the common mistake; the rail should carry ordinary spoken content.

What is the difference between Standard and Cinematic caption modes?

Standard mode uses the verbatim rail for most text with scarce embedded peaks, suited for explainers and voiceover. Cinematic mode drops the rail and makes everything embed-style, appropriate only for pure-cinematic asks.