tts

Synthesize speech from text using ElevenLabs, macOS say, or espeak.

650|137|Updated Jan 26, 2026
One-click install
npx skills add https://github.com/alsk1992/CloddsBot --skill tts-alsk1992
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: tts
Source: https://github.com/alsk1992/CloddsBot/tree/main/src/skills/bundled/tts
Command: npx skills add https://github.com/alsk1992/CloddsBot --skill tts-alsk1992

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

This Skill allows users to convert written text into spoken audio, making information more accessible and enabling voice-based interactions.

Core Features & Use Cases

  • Text-to-Speech Synthesis: Generate natural-sounding speech from text using various providers like ElevenLabs, macOS 'say', or 'espeak'.
  • Voice Customization: Select from a wide range of voices, adjust speech speed, pitch, and volume.
  • Use Case: Have an important alert or a long report read aloud to you while you're multitasking, or integrate voice output into your applications for a more interactive experience.

Quick Start

Use the tts skill to speak the phrase "Hello, world!" with the Rachel voice.

Frequently Asked Questions about tts

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert text to speech using different voice providers?

To convert text to speech, this Skill synthesizes spoken audio using multiple providers including ElevenLabs, macOS say, and espeak. You can generate natural-sounding speech by routing your text through any of these supported voice synthesis engines.

Can I adjust speed and pitch when generating audio from text?

Yes, you can adjust speech speed, pitch, and volume during text-to-speech synthesis. The Skill allows detailed voice customization, enabling you to select from a wide range of voices and fine-tune audio output parameters to suit your needs.

Does the text-to-speech synthesis support SSML for advanced control?

Yes, this text-to-speech synthesis supports SSML for advanced control over audio generation. By using SSML, you can precisely structure voice output, manage pronunciation, and direct the speech synthesis process beyond basic plain text inputs.

What is the best way to have a long report read aloud while multitasking?

The best way to have a long report read aloud is using this text-to-speech Skill to synthesize the text into audio. It generates natural-sounding speech from your documents and integrates with audio output devices for direct playback, enabling hands-free multitasking.

Do I need an ElevenLabs account to use this voice synthesis Skill?

You do not strictly need an ElevenLabs account to use this voice synthesis Skill, as it also supports macOS say and espeak for text-to-speech generation. However, utilizing ElevenLabs as a provider will require appropriate access credentials for its advanced voices.

Can I use this text-to-speech Skill for accessibility and voice-based interactions?

Yes, you can use this text-to-speech Skill for accessibility and voice-based interactions. It converts written text into spoken audio, making information more accessible and enabling interactive voice outputs directly through your connected audio output devices.