pr0ta-audio

Automate PR0TA audio production with ElevenLabs v3 TTS and Scribe V2 transcription.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/jeffamerican/pr0ta-plugin --skill pr0ta-audio
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: pr0ta-audio
Source: https://github.com/jeffamerican/pr0ta-plugin/tree/main/skills/pr0ta-audio
Command: npx skills add https://github.com/jeffamerican/pr0ta-plugin --skill pr0ta-audio

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

PR0TA audio generation and transcription streamlines end-to-end voice work: TTS, voice cloning, design prompts, STS, and transcription for storytelling in PR0TA productions.

Core Features & Use Cases

  • Text-to-speech (ElevenLabs v3) for narration and dialogue across PR0TA projects.
  • Voice cloning, design, and STS to adapt performances to new voices while preserving timing and emotion.
  • Speech-to-text transcription using Scribe V2 for word-level timing, speaker diarization, and event tagging to drive narration timelines.
  • Voice design and prompts to craft distinct voice personas for consistent on-brand narration.

Quick Start

Generate a short narration using ElevenLabs v3, then transcribe it with Scribe V2 to populate the narration timeline.

Frequently Asked Questions about pr0ta-audio

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate narration with ElevenLabs v3 and transcribe it for timeline assembly?

Generate narration using ElevenLabs v3 text-to-speech, then transcribe the output with Scribe V2 to extract word-level timing and speaker diarization for populating a narration timeline.

Can I clone a voice and transform speech while preserving the original timing?

Voice cloning and speech-to-speech transformation adapt performances to new voices while preserving original timing and emotion, enabling consistent character dialogue across different voice personas.

What is word-level transcription with diarization and when do I need it for audio workflows?

Word-level transcription with diarization tags individual speakers and timestamps each word, needed when aligning character dialogue or narration events to a precise production timeline.

Does this audio workflow support multilingual voice generation and model selection?

API-driven generation supports multilingual voice output and model selection, allowing you to choose specific TTS models and generate audio across multiple languages for diverse project requirements.

How do I create distinct voice personas for consistent character dialogue?

Use voice design prompts to craft distinct voice personas for consistent on-brand narration, defining specific vocal characteristics that remain uniform across generated dialogue and storytelling scenes.

What is the best way to build an end-to-end audio pipeline from generation to timeline?

Combine ElevenLabs v3 for text-to-speech generation, apply voice cloning or STS for performance adaptation, and run Scribe V2 transcription to assemble word-level timing data into a final narration timeline.