Audio Generation

Generate MP3 audio from text using ElevenLabs, OpenAI TTS, or Google Text-to-Speech.

24|26|Updated Feb 14, 2026
One-click install
npx skills add https://github.com/hasna/skills --skill audio-generation-hasna
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: Audio Generation
Source: https://github.com/hasna/skills/tree/main/skills/skill-audio
Command: npx skills add https://github.com/hasna/skills --skill audio-generation-hasna

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

This Skill allows users to convert text into natural-sounding speech using advanced AI models, eliminating the need for manual voice recording or expensive studio time.

Core Features & Use Cases

  • Multi-Provider Support: Utilizes ElevenLabs, OpenAI TTS, and Google Text-to-Speech for diverse voice options and quality.
  • Voice Customization: Select from a wide array of voices, languages, and adjust speech speed for tailored audio output.
  • Use Case: Generate podcast intros, audiobook narration, or voiceovers for video content quickly and efficiently.

Quick Start

Generate an MP3 audio file from the text "Hello, world!" using the OpenAI TTS service with the 'nova' voice.

Frequently Asked Questions about Audio Generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert text to speech using AI for audio generation?

AI text-to-speech generation converts written text into natural-sounding speech using APIs like ElevenLabs, OpenAI TTS, and Google Text-to-Speech, eliminating manual voice recording. It outputs high-quality MP3 audio files suitable for content creation and accessibility.

What API keys do I need for AI text-to-speech generation?

AI text-to-speech generation requires valid API keys for ElevenLabs, OpenAI, or Google Cloud. You must configure at least one of these provider keys in your environment to authenticate requests and generate audio output successfully.

Can I customize voice selection and speech speed for text-to-speech output?

Text-to-speech voice customization supports selecting from a wide array of voices, multiple languages, and adjustable speech speed. This allows tailored audio output using ElevenLabs, OpenAI TTS, or Google Text-to-Speech APIs for specific narration needs.

What is the best way to generate podcast intros or audiobook narration from text?

The best way to generate podcast intros or audiobook narration is using AI text-to-speech APIs like ElevenLabs and OpenAI TTS. They produce high-quality, natural-sounding audio quickly, bypassing expensive studio time and manual recording processes.

Does text-to-speech generation work with multiple languages for automated narration?

AI text-to-speech generation supports multiple languages for automated narration across ElevenLabs, OpenAI TTS, and Google Text-to-Speech. This multilingual capability enables diverse voice options for accessible content creation and global audience reach.