ai-media

Generate images, videos, and voiceovers via inference.sh CLI and Remotion.

13|3|Updated Mar 2, 2026
One-click install
npx skills add https://github.com/phrazzld/agent-skills --skill ai-media
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: ai-media
Source: https://github.com/phrazzld/agent-skills/tree/main/packs/growth/ai-media
Command: npx skills add https://github.com/phrazzld/agent-skills --skill ai-media

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill streamlines the creation of various media assets, from static images and dynamic videos to voiceovers and programmatic video content, leveraging a vast array of AI models.

Core Features & Use Cases

  • Multi-modal Generation: Supports text-to-image, text-to-video, image-to-video, avatar generation, and voiceover synthesis.
  • Programmatic Video: Enables the creation of videos from React components using Remotion.
  • Asset Pipelines: Includes pipelines for generating logos, icons, and demo videos.
  • Use Case: A marketing team needs a short promotional video for a new product. They can use this Skill to generate a script, create a voiceover, and then render a video with relevant visuals, all through a single interface.

Quick Start

Use the ai-media skill to generate an image of a cat astronaut in space.

Frequently Asked Questions about ai-media

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate images and videos from text prompts using AI models?

Generating images and videos from text prompts is facilitated by a CLI interface connecting to over 150 models. This enables text-to-image and text-to-video generation by routing inputs through providers like Google, OpenAI, and xAI.

Can I create programmatic videos using Remotion and AI voiceovers?

Creating programmatic videos using Remotion and AI voiceovers is fully supported. The system renders videos from React components via Remotion and synthesizes voiceovers using ElevenLabs, allowing you to combine dynamic visuals with generated audio in a single workflow.

Does this media generation pipeline support image-to-video and avatar lipsync?

This media generation pipeline supports image-to-video and avatar lipsync. It leverages a CLI tool to access over 150 models, enabling you to animate static images and generate lip-synced avatars alongside standard text-to-image and voiceover synthesis tasks.

What is the best way to build an automated marketing asset pipeline?

Building an automated marketing asset pipeline is best achieved using a multi-modal generation interface that handles scripts, voiceovers, and video rendering. This allows you to generate logos, icons, and demo videos through a single CLI supporting diverse model providers.

How do I generate voiceovers with ElevenLabs for my video assets?

Generating voiceovers with ElevenLabs for your video assets is done through the media generation CLI. This integrates voiceover synthesis directly into your asset pipeline, allowing you to create audio tracks for programmatic videos or standalone demo content.