speech-synthesis

Convert text to speech using Microsoft Edge neural voices and generate MP3 audio.

12|2|Updated Apr 21, 2026
One-click install
npx skills add https://github.com/haomingz/kimi-skills --skill speech-synthesis
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: speech-synthesis
Source: https://github.com/haomingz/kimi-skills/tree/main/skills/speech-synthesis
Command: npx skills add https://github.com/haomingz/kimi-skills --skill speech-synthesis

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires node-edge-tts, commander, and includes scripts (resource) and references (resource) components.

What problem does it solve?

将文本转换为高质量语音,支持多语言、多种音色,可调节语速、音调和音量,并生成字幕,输出MP3音频文件。当你需要朗读文本(比如文章、消息)、为视频或演示文稿生成配音,或直接提到“TTS”、“文字转语音”、“配音”、“朗读”时,就会使用这个技能。

Core Features & Use Cases

  • 支持多语言、多音色、可调节语速、音调和音量,并输出 MP3 音频。
  • 生成字幕,便于视频和内容的无障碍呈现。
  • 常见用例包括为文章、消息、视频演示生成配音,以及将文本转换为可分享的音频文件。

Quick Start

输入要转换为语音的文本,并可选择音色、语言和输出格式,即可生成音频文件。

Frequently Asked Questions about speech-synthesis

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert text to natural speech in multiple languages?

This text-to-speech tool converts text to natural speech by using Microsoft Edge neural voices to generate MP3 audio, supporting multiple languages with adjustable rate, pitch, and volume.

Can I generate subtitles along with text-to-speech audio?

Yes, this text-to-speech process supports subtitle generation alongside MP3 audio output, facilitating accessible video and content presentation.

Does this text-to-speech tool support adjusting voice rate, pitch, and volume?

Yes, the voice synthesis process allows you to adjust rate, pitch, and volume for Microsoft Edge neural voices before outputting the final MP3 audio file.

Do I need node-edge-tts to generate MP3 audio from text?

Yes, you need the node-edge-tts dependency installed to execute this text-to-speech functionality, as it provides the Microsoft Edge neural voices required to synthesize MP3 audio.

What is the best way to generate voiceovers for video presentations?

The best way to generate voiceovers for presentations is to use this text-to-speech capability, leveraging Microsoft Edge neural voices to customize voice, rate, pitch, and output MP3 files with subtitles.