media

Generate images, videos, transcripts, and speech via harness CLI media commands.

28|2|Updated Feb 5, 2026
One-click install
npx skills add https://github.com/thompson0012/agents-stack --skill media-thompson0012
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: media
Source: https://github.com/thompson0012/agents-stack/tree/main/skills-optional/media
Command: npx skills add https://github.com/thompson0012/agents-stack --skill media-thompson0012

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill enables you to generate media content and transcribe media assets using the harness CLI commands, reducing manual steps for images, video, speech, and audio tasks.

Core Features & Use Cases

  • Image and video generation with asi-generate-image and asi-generate-video for quick visual assets.
  • Speech and audio transcription using asi-transcribe-audio for captions, transcripts, and diarization.
  • Text-to-speech options via asi-text-to-speech to generate voiceovers with multiple voices and delivery controls.
  • Use Case: Create a short product demo video by generating visuals, narrating with TTS, and exporting transcripts for captions.

Quick Start

Generate a media asset by issuing a single media command with a prompt.

Frequently Asked Questions about media

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate images and video from text prompts using a CLI?

You can generate images and video from text prompts by issuing single harness CLI media commands like asi-generate-image and asi-generate-video to quickly produce visual assets.

Can I transcribe audio to text and get diarization for video captions?

Yes, you can transcribe audio to text with diarization for captions and transcripts using the asi-transcribe-audio command provided by the harness media CLI.

What is the best way to create voiceovers from text for a product demo?

The best way to create voiceovers from text is using the asi-text-to-speech command, which generates speech with multiple voices and delivery controls for media outputs.

Does text-to-speech synthesis support multiple voices for audio generation?

Yes, text-to-speech synthesis supports multiple voices and delivery controls, allowing you to generate customized voiceovers and audio assets directly from text prompts.

Are there limitations or safety guidelines for AI media generation via CLI?

AI media generation via CLI requires adherence to safety guidelines for content generation, meaning you must ensure prompts and outputs comply with content policies when generating images, video, or speech.