sag

Generate natural-sounding speech from text and manage local playback via the sag CLI.

1|Updated May 12, 2026
One-click install
npx skills add https://github.com/estebanrfp/gos --skill sag-estebanrfp
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sag
Source: https://github.com/estebanrfp/gos/tree/main/skills/sag
Command: npx skills add https://github.com/estebanrfp/gos --skill sag-estebanrfp

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

sag removes the friction of turning written text into polished spoken audio, making it easy to produce voice replies, narration, and quick playback without juggling multiple tools.

Core Features & Use Cases

  • Natural speech generation for chat replies, scripts, and short announcements.
  • Voice selection and style control for expressive delivery, including model choices, normalization, and pronunciation tuning.
  • Practical workflows for assistants, creators, and support agents who need fast, consistent ElevenLabs audio output.

Quick Start

Use the sag skill to turn your text into an ElevenLabs voice reply and tell me which voice you want if you already have one in mind.

Frequently Asked Questions about sag

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I convert text to natural speech using ElevenLabs voice synthesis?

Text-to-speech conversion with ElevenLabs voice synthesis is handled by generating natural-sounding audio from your text and managing local playback. You can produce voice replies, narration, and expressive spoken content by selecting a specific voice and applying normalization options.

Do I need an ElevenLabs API key to generate voice audio from text?

Yes, an ElevenLabs API key is required to generate voice audio from text. The voice synthesis process depends on this key to authenticate requests, select your preferred voice, and deliver compatible audio output through the local playback manager.

Can I control voice style and pronunciation for text-to-speech narration?

Yes, you can control voice style and pronunciation for text-to-speech narration. The generation process supports voice selection, model choices, normalization options, and pronunciation tuning to ensure expressive delivery for scripts, chat replies, and announcements.

What is the best way to add spoken responses to an assistant workflow?

The best way to add spoken responses to an assistant workflow is using a dedicated text-to-speech skill that normalizes text and manages audio playback. This provides fast, consistent ElevenLabs voice output for support agents and interactive chat use cases.

Does text-to-speech voice generation work for short announcements and scripts?

Text-to-speech voice generation works effectively for short announcements, scripts, and chat replies. It transforms written text into polished spoken audio, removing the friction of juggling multiple tools for quick narration and media playback.

Why is my ElevenLabs voice output not playing back locally?

ElevenLabs voice output requires the sag CLI for compatible local audio delivery. If playback is not working, ensure the CLI is properly configured and that your voice selection and API key are correctly specified in the text-to-speech generation parameters.