voice

Generate text-to-speech audio and sound effects synchronized with video timelines.

907|122|Updated Jul 15, 2026
One-click install
npx skills add https://github.com/0xsline/OpenChatCut --skill voice-0xsline
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: voice
Source: https://github.com/0xsline/OpenChatCut/tree/main/src/agent/skills/voice
Command: npx skills add https://github.com/0xsline/OpenChatCut --skill voice-0xsline

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This skill solves the challenge of creating high-quality, synchronized audio assets for video projects without needing external recording equipment or complex audio editing software.

Core Features & Use Cases

  • Text-to-Speech (TTS) Generation: Convert scripts into natural-sounding narration using multiple providers like Doubao, ElevenLabs, and MiniMax.
  • Video-Audio Synchronization: Automatically align voiceover timing with visual events, scene changes, or on-screen actions.
  • Custom Sound Effects: Generate specific, context-aware sound effects when library assets are insufficient.
  • Use Case: Quickly generate a professional voiceover for a product demo video and ensure the narration perfectly matches the timing of the on-screen feature highlights.

Quick Start

Use the voice skill to generate a professional English voiceover for my video using the ElevenLabs provider and the Peter voice preset.

Frequently Asked Questions about voice

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate a voiceover for video editing without external recording equipment?

To generate a voiceover for video editing without external recording, use text-to-speech to convert scripts into natural narration, then synchronize the audio segments with your project timeline and visual anchors.

Can I use text-to-speech providers like ElevenLabs and Doubao for video narration?

Yes, you can use text-to-speech providers like ElevenLabs and Doubao for video narration. The skill supports multi-provider synthesis, allowing you to select specific voice presets to generate high-quality audio for your production workflow.

How does audio sync work when adding a voiceover to a video timeline?

Audio sync aligns generated text-to-speech voiceover timing with your video timeline data. It ensures audio segments match on-screen visual events, scene changes, and narrative beats by integrating with project timeline data.

What is the best way to create custom sound effects when library assets are insufficient?

The best way to create custom sound effects when library assets are insufficient is to generate context-aware audio tailored to your specific video. This allows you to produce specific sound effects that match on-screen actions without complex audio editing software.

Does generating TTS audio require integration with project timeline data?

Yes, generating TTS audio for video production requires integration with project timeline data. This ensures precise timing and synchronization capabilities, allowing the generated audio segments to align correctly with visual media anchors.