What problem does it solve? Choosing and wiring the right provider for image generation, TTS, and speech-to-text is fragmented across dozens of vendors with different cost, latency, and privacy trade-offs. This Skill consolidates provider selection, cost control, and compliance rules into one operational guide. ## Core Features & Use Cases - Provider Selection Matrices: Decision tables for image generation (Imagen, gpt-image, SDXL, Flux), TTS (ElevenLabs, Deepgram, Coqui), and STT (Whisper, AssemblyAI, Deepgram) based on fidelity, latency, and hosting needs. - Cost & Latency Controls: Guidance on caching TTS by voice and text hash, VAD pre-processing with silero-vad, per-feature daily caps, and streaming for latency-critical paths. - Compliance Guardrails: Hard rules for voice-cloning consent, DPA requirements before sending customer audio to third parties, retention limits, and C2PA provenance tagging. - Use Case: A team building a voice-enabled support agent uses this Skill to pick Deepgram for streaming STT, ElevenLabs Flash for low-latency TTS, and configure cost caps plus PII redaction before storing transcripts. ## Quick Start Ask the agent to pick a speech-to-text provider and wire it up for transcribing support calls with diarization and PII redaction.