sag

Generate ElevenLabs speech from text and play it locally via the sag CLI.

5|Updated Jan 31, 2026
One-click install
npx skills add https://github.com/kcns008/clusterclaw --skill sag-kcns008
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: sag
Source: https://github.com/kcns008/clusterclaw/tree/main/skills/sag
Command: npx skills add https://github.com/kcns008/clusterclaw --skill sag-kcns008

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

sag removes the friction of generating natural-sounding speech from text, so you can create voice replies, narrations, and spoken prompts without leaving the terminal.

Core Features & Use Cases

  • ElevenLabs TTS Playback: Produce expressive speech with local playback for quick conversational responses and short-form narration.
  • Voice Control: Choose voices, set defaults, and keep speaker selection consistent for repeated use.
  • Pronunciation Guidance: Improve difficult words, numbers, URLs, and pacing with normalization rules, pause markers, and expressive audio tags.
  • Use Case: A developer can turn a status update, explanation, or “voice reply” request into an audio file and play it immediately on a Mac-style workflow.

Quick Start

Use sag to turn your text into spoken audio with your preferred voice and play it locally.

Frequently Asked Questions about sag

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate ElevenLabs text-to-speech audio and play it locally from the terminal?

To generate ElevenLabs text-to-speech audio locally, you can use a terminal-first voice workflow that converts text into spoken media files and plays them immediately. This requires the sag CLI and an ElevenLabs API key.

Can I control pronunciation and pacing for text-to-speech narration?

Yes, you can control pronunciation and pacing in text-to-speech narration by applying normalization rules, pause markers, and expressive audio tags to improve difficult words, numbers, and URLs.

What's the best way to keep voice selection consistent for repeated text-to-speech output?

The best way to maintain consistent voice selection for text-to-speech output is to set a default voice ID. This ensures the same speaker is used across repeated audio playback and narrations.

Do I need an ElevenLabs API key to generate expressive speech from text?

Yes, an ElevenLabs API key is required to generate expressive speech from text. The voice workflow relies on this key to produce natural-sounding audio files for local playback.

Does this text-to-speech workflow support narrating long-form text content?

Yes, the text-to-speech workflow supports both short conversational replies and long-form audio output, allowing you to narrate extensive explanations or status updates as spoken media files.

What are the limitations of terminal-based voice output for text narration?

Terminal-based voice output for text narration requires a Mac-style local playback environment and optional voice ID settings. Without proper pronunciation rules, complex words or URLs may not sound natural.