jarvis-voice

Generate local metallic text-to-speech with sherpa-onnx and customizable effects.

Updated Feb 17, 2026
One-click install
npx skills add https://github.com/Qcasares/saas-app --skill jarvis-voice-qcasares
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: jarvis-voice
Source: https://github.com/Qcasares/saas-app/tree/main/skills/jarvis-voice
Command: npx skills add https://github.com/Qcasares/saas-app --skill jarvis-voice-qcasares

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires sherpa-onnx, ffmpeg, aplay, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill provides a local, metallic text-to-speech (TTS) voice for AI agents, enhancing user engagement and accessibility through visual transcripts and voice effects.

Core Features & Use Cases

  • Metallic Voice: Local speech synthesis with sherpa-onnx, avoiding cloud latency.
  • Customizable Effects: Apply effects like flanger, echo, and pitch shift for unique voices.
  • Visual Transcripts: Differentiate spoken text visually in webchat with purple italic styling.
  • Fast Playback: Enable 2x speed for quick communication.
  • Use Case: Improve the user experience for AI-powered chatbots and assistants by adding a robotic, distinctive voice that can be customized with various effects.

Quick Start

Use the jarvis-voice skill to generate a speech output for the text 'Hello, I am your AI assistant.'.

Frequently Asked Questions about jarvis-voice

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I add local text-to-speech to an AI agent without cloud APIs?

Local text-to-speech for AI agents can be generated using sherpa-onnx to synthesize speech without relying on cloud APIs. This approach processes audio locally to avoid cloud latency while providing customizable voice effects and visual transcripts.

What voice effects can I apply to local TTS output for chatbots?

Voice effects for local TTS output include flanger, echo, and pitch shift to create unique metallic voices. These effects allow you to customize speech synthesis, enhancing user engagement for AI-powered chatbots and assistants.

Does local TTS processing require sherpa-onnx and ffmpeg?

Local TTS processing requires sherpa-onnx, ffmpeg, and aplay to function properly without cloud dependencies. These tools handle speech synthesis, audio processing, and playback respectively, ensuring the local metallic voice generation works as intended.

Can I display visual transcripts alongside speech output in webchat?

Visual transcripts can be displayed alongside speech output in webchat using purple italic styling to differentiate spoken text. This feature improves accessibility and user engagement by providing a clear visual representation of the synthesized speech.

How do I enable fast playback for AI assistant voice output?

Fast playback for AI assistant voice output can be enabled by applying a 2x speed setting to the synthesized speech. This allows for quick communication, reducing the time users spend listening to responses from the local metallic TTS system.