audio

Generate text-to-speech audio and clone voices with specified output files.

1|Updated Jun 4, 2026
One-click install
npx skills add https://github.com/hung-phan/ml-skills --skill audio-hung-phan
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: audio
Source: https://github.com/hung-phan/ml-skills/tree/main/skills/ml-review/references/ml-architectures/audio
Command: npx skills add https://github.com/hung-phan/ml-skills --skill audio-hung-phan

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides a comprehensive solution for audio generation, voice cloning, and music creation, addressing the challenges of selecting the right models, understanding architecture nuances, and creating high-quality audio outputs.

Core Features & Use Cases

  • Audio Generation: Text-to-speech (TTS), music generation, and sound effect creation.
  • Voice Cloning: Replicating a voice for different applications.
  • Use Case: Imagine you need a custom voice for a new voice assistant. Use this Skill to clone a specific voice and customize it to suit your needs.

Quick Start

To generate a TTS from the provided text, use the audio skill with the 'generate' command, specifying the desired output file name.

Frequently Asked Questions about audio

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate text-to-speech audio from a text file?

To generate text-to-speech audio, use the generate command and specify the desired output file name. This Skill supports TTS creation alongside music generation and sound effect outputs.

Can I clone a specific voice for a custom voice assistant?

Yes, voice cloning is supported to replicate a specific voice for different applications. You can clone and customize a voice to suit custom voice assistant needs using various synthesis model architectures.

What is the best way to choose the right models for audio generation?

Selecting the right models for audio generation requires understanding architecture nuances and codec functions. This Skill helps navigate different speech and music synthesis models to create high-quality audio outputs.

Do I need to understand audio codecs to use voice synthesis techniques?

Understanding audio file formats and codec functions is required for effective voice synthesis. This knowledge helps you grasp the nuances of different speech synthesis models for proper audio manipulation.

Does this Skill support both music generation and sound effect creation?

Music generation and sound effect creation are both supported as core audio generation features. The Skill handles these alongside text-to-speech and voice cloning using various model architectures.