ai-captions

Integrate Deepgram transcription via LiveKit data channels for live video room captions.

3|Updated Feb 14, 2026
One-click install
npx skills add https://github.com/mattwoodco/skills --skill ai-captions
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ai-captions
Source: https://github.com/mattwoodco/skills/tree/main/skills/ai-captions
Command: npx skills add https://github.com/mattwoodco/skills --skill ai-captions

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides real-time, AI-powered closed captions for video conferencing, making content accessible and understandable for everyone.

Core Features & Use Cases

  • Live Transcription: Displays spoken words as text in real-time using Deepgram.
  • Speaker Labels: Identifies who is speaking.
  • Customization: Allows users to adjust font size, position, and background opacity.
  • Use Case: Enhance accessibility in online meetings or live streams by providing instant captions for hearing-impaired users or those in noisy environments.

Quick Start

Integrate the ai-captions skill into your video room component to display live captions.

Frequently Asked Questions about ai-captions

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I add live closed captions to a video room?

Live closed captions for video rooms are enabled by integrating Deepgram transcription via LiveKit data channels. This processes spoken audio in real-time and displays the resulting text as an overlay in the video interface.

Does LiveKit data channels support real-time Deepgram transcription?

Yes, LiveKit data channels support real-time Deepgram transcription. This combination transports the AI-generated speech-to-text data from the audio stream directly into the video UI components for display.

Can I customize the floating caption overlay for speaker labels and font size?

Yes, the floating caption overlay is fully configurable. You can adjust the caption font size, screen position, background opacity, and display speaker labels to identify who is currently speaking.

What is the best way to transcribe spoken words as text in real-time during online meetings?

Using Deepgram transcription integrated with LiveKit data channels is an effective way to display spoken words as text in real-time. It provides instant transcription and speaker labeling directly within the video conferencing UI.

Do I need Deepgram transcription to display live captions for hearing-impaired users?

Yes, Deepgram transcription is required to process the audio and generate the live closed captions. This provides real-time accessibility for hearing-impaired users or those in noisy environments.