gemini-image-generator

Generate and edit images from text prompts via Google's Gemini API.

12|Updated Dec 13, 2025
One-click install
npx skills add https://github.com/ckorhonen/claude-skills --skill gemini-image-generator-ckorhonen
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: gemini-image-generator
Source: https://github.com/ckorhonen/claude-skills/tree/main/skills/gemini-image-generator
Command: npx skills add https://github.com/ckorhonen/claude-skills --skill gemini-image-generator-ckorhonen

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires google-genai, and includes scripts (resource) components.

What problem does it solve?

Generate high-quality images from text prompts, edit existing images with prompts, and combine reference images to rapidly produce visual assets for apps, marketing, or games.

Core Features & Use Cases

  • Text-to-image generation from prompts for UI assets, marketing visuals, game art, and concept illustrations.
  • Image editing and refinement by prompting changes to an existing image.
  • Multi-image reference input to guide style and composition across multiple assets.
  • Supports Gemini 2.5 Flash (fast iterations) and Gemini 3 Pro (high quality) with up to 4K output, suitable for professional workflows.

Quick Start

Provide a text prompt and optional reference images, then run the Gemini Image Generator to create your image.

Frequently Asked Questions about gemini-image-generator

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate images from text prompts using the Gemini API?

To generate images from text prompts using the Gemini API, provide a prompt and optional reference images to the skill. It outputs high-quality visual assets for UI, marketing, or game art using Google's gemini-2.5-flash-image or gemini-3-pro-image-preview models.

Can I edit existing images and combine multiple reference images with Gemini?

Yes, you can edit existing images and combine multiple reference images with Gemini by prompting changes to guide style and composition. The skill processes these inputs to refine visuals and maintain consistency across multiple generated assets.

Do I need a specific API key to use Gemini for text-to-image generation?

Yes, you need a GEMINI_API_KEY to use Gemini for text-to-image generation. The skill uses this key to authenticate requests to Google's gemini-2.5-flash-image and gemini-3-pro-image-preview models for creating and editing images.

What is the best way to create high-resolution marketing visuals with AI?

The best way to create high-resolution marketing visuals with AI is using Gemini 3 Pro, which supports up to 4K output. This skill handles text-to-image generation and image editing suitable for professional design workflows.

Does the Gemini image generator support configurable aspect ratios and output sizes?

Yes, the Gemini image generator supports configurable aspect ratios and output sizes. It includes robust input validation and error handling to ensure generated UI assets and concept previews meet specific design requirements.

When should I use Gemini 2.5 Flash versus Gemini 3 Pro for image generation?

Use Gemini 2.5 Flash for fast iterations during image generation, and use Gemini 3 Pro for high-quality outputs up to 4K. The skill supports both models, allowing you to choose based on whether you need speed or professional visual fidelity.