alicloud-ai-audio-cosyvoice-voice-clone

Enroll a custom voice from reference audio using Alibaba Cloud CosyVoice models.

396|34|Updated Jan 31, 2026
One-click install
npx skills add https://github.com/cinience/alicloud-skills --skill alicloud-ai-audio-cosyvoice-voice-clone
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: alicloud-ai-audio-cosyvoice-voice-clone
Source: https://github.com/cinience/alicloud-skills/tree/main/skills/ai/audio/alicloud-ai-audio-cosyvoice-voice-clone
Command: npx skills add https://github.com/cinience/alicloud-skills --skill alicloud-ai-audio-cosyvoice-voice-clone

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill automates the creation of custom cloned voices from reference audio, enabling personalized speech synthesis.

Core Features & Use Cases

  • Voice Cloning: Create unique voice profiles using Alibaba Cloud's CosyVoice models.
  • TTS Integration: Reuses the cloned voice_id for subsequent text-to-speech synthesis.
  • Use Case: Generate a branded voice for your AI assistant or customer service bot using a specific actor's voice sample.

Quick Start

Use the alicloud-ai-audio-cosyvoice-voice-clone skill to clone a voice using the provided audio sample and prefix.

Frequently Asked Questions about alicloud-ai-audio-cosyvoice-voice-clone

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I clone a voice using a reference audio sample?

To clone a voice, you provide a reference audio URL to enroll a custom voice profile. The system processes this sample to generate a reusable voice ID for subsequent text-to-speech synthesis.

How does voice cloning work with Alibaba Cloud CosyVoice?

Voice cloning with CosyVoice enrolls a custom voice using Alibaba Cloud Model Studio models v3.5-plus or v3.5-flash. It generates a reusable voice ID from a provided reference audio URL for personalized speech synthesis.

Can I use a cloned voice ID for text-to-speech synthesis?

Yes, you can reuse the cloned voice ID for subsequent text-to-speech synthesis. This allows you to generate speech output that mimics the specific actor's voice sample used during enrollment.

What audio formats or inputs are needed to create a custom voice profile?

You need a reference audio URL to create a custom voice profile. The system supports various language hints and optional preprocessing to ensure accurate voice cloning and synthesis.

Are there different models available for CosyVoice voice cloning?

Yes, CosyVoice offers v3.5-plus and v3.5-flash models for voice cloning. These models process your reference audio URL to generate a unique voice ID for text-to-speech applications.