ai-voice-cloning

Generate natural-sounding speech from text using multiple TTS models.

688|95|Updated Jan 31, 2026
One-click install
npx skills add https://github.com/inference-sh/skills --skill ai-voice-cloning-inference-sh
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: ai-voice-cloning
Source: https://github.com/inference-sh/skills/tree/main/tools/audio/ai-voice-cloning
Command: npx skills add https://github.com/inference-sh/skills --skill ai-voice-cloning-inference-sh

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill automates the creation of natural-sounding AI voices for various audio needs, eliminating the need for expensive recording equipment or voice actors.

Core Features & Use Cases

  • Text-to-Speech (TTS): Convert written text into spoken audio using a variety of models and voices.
  • Voice Cloning: Synthesize speech in specific voices (requires appropriate permissions and models).
  • Customization: Control voice style, emotion, accent, and speaking speed for tailored audio output.
  • Use Case: Generate a professional audiobook narration, create engaging voiceovers for marketing videos, or produce realistic dialogue for AI-powered characters in games.

Quick Start

Use the ai-voice-cloning skill to generate speech from the text "Hello, this is a test of the AI voice cloning skill." using the 'af_sarah' voice.

Frequently Asked Questions about ai-voice-cloning

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert text to speech for a video narration?▼

This Skill performs text-to-speech synthesis for video narration using models like Kokoro TTS, DIA, and Chatterbox. It generates natural-sounding AI voices, automating audio creation and eliminating the need for expensive recording equipment or voice actors.

Can I control the emotion and accent of AI generated voices?▼

Yes, you can control the emotion, accent, and speaking speed of AI generated voices. The Skill supports various voice styles and customizations, allowing you to produce tailored audio output for professional audiobook narration, podcasts, or realistic game character dialogue.

What is the best way to generate audiobook narration without a voice actor?▼

The best way to generate audiobook narration without a voice actor is using this Skill's text-to-speech synthesis. It leverages multiple TTS models to create natural-sounding AI voices, providing a cost-effective alternative to hiring voice actors or purchasing recording equipment.

Does voice cloning work for creating dialogue in AI-powered game characters?▼

Yes, voice cloning works for creating realistic dialogue in AI-powered game characters. The Skill synthesizes speech in specific voices using models like DIA and Chatterbox, allowing you to generate customized voiceovers with various emotions and accents for interactive applications.

How do I start text to speech synthesis using the inference CLI?▼

To start text to speech synthesis using the inference CLI, provide your text and select a voice, such as 'af_sarah'. The Skill leverages the inference.sh CLI for deterministic task execution, quickly converting your written text into spoken audio output.