media-generation

Generate images, videos, 3D models, and audio from text descriptions.

Updated Jun 30, 2026
One-click install
npx skills add https://github.com/not-subhu/screech --skill media-generation-not-subhu
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: media-generation
Source: https://github.com/not-subhu/screech/tree/main/.local/skills/media-generation
Command: npx skills add https://github.com/not-subhu/screech --skill media-generation-not-subhu

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) and assets (resource) components.

What problem does it solve?

This Skill solves the problem of manually creating various forms of media, including images, videos, 3D models, and audio. It enables the automation of content creation processes for efficiency and customization.

Core Features & Use Cases

  • Image Generation: Create custom images from text descriptions, with options for removing backgrounds and resolution settings.
  • Video Generation: Generate short video clips from detailed text descriptions, with control over aspect ratio, resolution, and duration.
  • Audio Generation: Create music, sound effects, and text-to-speech audio from prompts.
  • 3D Model Generation: Generate static 3D models based on text descriptions, with quality settings for output.
  • Use Case: Imagine you need to create a promotional video for a new product. Use this Skill to generate images of the product, its promotional video, a 3D model for the product display, and background music for the video.

Quick Start

Generate a high-quality image of a 'futuristic cityscape at night' and save it to 'assets/cityscape.jpg'.

Frequently Asked Questions about media-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate images and videos from text descriptions without manual editing?

You can generate images and videos from text descriptions by using Python and machine learning libraries to automate media creation. This process supports custom resolution, aspect ratio, and duration settings without requiring manual design or editing software.

Can I create 3D models and audio from text prompts for a promotional video?

Yes, you can create static 3D models, music, sound effects, and text-to-speech audio from text prompts. This allows you to generate all necessary media assets for a promotional video, including product displays and background music.

Does media generation support image editing like background removal and resolution settings?

Image generation supports background removal and custom resolution settings. These options are available directly through text-based automation, allowing you to refine the output without manual image editing tools.

What is the best way to automate content creation for multiple media types like audio and video?

The best way to automate content creation for multiple media types is to use a script-based approach with machine learning libraries. This handles image, video, 3D model, and audio synthesis simultaneously from text inputs.

Are there limitations on video generation duration and aspect ratio when using text prompts?

Video generation allows control over aspect ratio, resolution, and duration, but is designed for short video clips based on detailed text descriptions. Outputs are automated via Python without manual editing.

Do I need Python and machine learning libraries to generate 3D models from text?

Yes, Python and machine learning libraries are required to generate static 3D models from text descriptions. This setup automates the model creation process and allows you to configure quality settings for the output.