text-to-speech

Generate spoken audio from text using ElevenLabs models and voice settings.

1|Updated Apr 19, 2026
One-click install
npx skills add https://github.com/arthtyagi/onloop --skill text-to-speech-arthtyagi
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: text-to-speech
Source: https://github.com/arthtyagi/onloop/tree/main/.agents/skills/text-to-speech
Command: npx skills add https://github.com/arthtyagi/onloop --skill text-to-speech-arthtyagi

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill removes the friction of producing high-quality spoken audio from written text, making it easy to create narration, voiceovers, and spoken content without manual recording.

Core Features & Use Cases

  • Natural Speech Generation: Convert text into realistic speech with ElevenLabs voices and models optimized for quality or latency.
  • Voice App and Media Support: Build voice-enabled apps, generate audiobook-style narration, and produce audio for podcasts, demos, and product experiences.
  • Streaming and Fine Control: Use streaming output for low-latency playback and adjust voice settings like stability, similarity, style, and speed for different delivery styles.

Quick Start

Use the text-to-speech skill to turn the provided script into a natural ElevenLabs voiceover and save the result as an MP3.

Frequently Asked Questions about text-to-speech

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert text to speech for a voiceover?

You can generate natural voiceovers by processing your written text with ElevenLabs speech synthesis models. This Skill converts scripts into realistic spoken audio, allowing you to produce narration and save the result as an MP3 file.

Do I need an ElevenLabs API key to generate audio?

Yes, you need an ElevenLabs API key to use this Skill. The text-to-speech generation relies directly on ElevenLabs models and voice settings, requiring your own valid API key to authenticate requests and synthesize spoken audio.

Can I adjust voice settings like stability and speed for text to speech?

Yes, you can fine-tune voice settings including stability, similarity, style, and speed. These parameters allow you to adjust the ElevenLabs voice delivery to match different narration styles and specific audio production requirements.

Does this support streaming audio for low-latency playback?

Yes, the Skill supports streaming audio output for low-latency playback. By using ElevenLabs models optimized for latency, you can generate real-time audio streams suitable for voice-enabled apps and interactive product experiences.

What is the best way to create multilingual speech from text?

The best way to create multilingual speech is by using ElevenLabs text-to-speech models. This Skill supports multilingual speech synthesis, enabling you to convert written scripts into realistic spoken audio across various languages for global voiceovers.

Can I use this for audiobook narration and podcast audio generation?

Yes, you can use this Skill for audiobook narration and podcast audio generation. It converts written text into spoken content optimized for media production, eliminating the need for manual recording while maintaining natural delivery.