gemini-imagegen

Generate and edit images from text prompts and reference images via the Gemini API.

1|Updated Jun 28, 2026
One-click install
npx skills add https://github.com/whmathews15/DEX-Personal-Operating-System --skill gemini-imagegen-whmathews15
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: gemini-imagegen
Source: https://github.com/whmathews15/DEX-Personal-Operating-System/tree/main/.claude/plugins/compound-engineering/skills/gemini-imagegen
Command: npx skills add https://github.com/whmathews15/DEX-Personal-Operating-System --skill gemini-imagegen-whmathews15

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires google-genai, Pillow, and includes scripts (resource) components.

What problem does it solve?

It removes the friction of moving from a creative idea to a usable image by turning prompts, reference photos, and iterative feedback into generated or edited visuals.

Core Features & Use Cases

  • Text-to-Image Generation: Create new visuals from a written prompt with control over aspect ratio and resolution.
  • Image Editing: Modify an existing image with conversational instructions, such as changing colors, adding objects, or shifting style.
  • Multi-Image Composition: Combine multiple reference images into one result for scene building, group shots, and composite concepts.
  • Use Case: A designer can generate a logo draft, refine it in chat, then compose a final mockup from several reference images without leaving the workflow.

Quick Start

Use the gemini-imagegen skill to create a polished product mockup from a short prompt and save the generated image to a file.

Frequently Asked Questions about gemini-imagegen

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate and edit images using the Gemini API?

You can generate and edit images using the Gemini API by submitting text prompts and reference images to the gemini-imagegen workflow. It supports conversational visual refinement, allowing you to modify colors, add objects, or shift styles without leaving your environment.

Can I combine multiple reference images into a single generated mockup?

Yes, you can compose multiple reference images into a single generated result. The Skill supports multi-image composition with up to 14 reference images, enabling you to build complex scenes, group shots, and composite product mockups from existing visuals.

Do I need a specific API key to perform text-to-image generation?

Yes, you need a configured GEMINI_API_KEY to perform text-to-image generation. The Skill requires the google-genai dependency to authenticate and interact with the Gemini API for creating new visuals from written prompts.

What's the best way to refine a logo draft through conversational image editing?

The best way to refine a logo draft is through multi-turn visual refinement. You can generate an initial design from a text prompt and then apply iterative, conversational instructions to modify specific elements like colors and composition until the final logo is achieved.

Does this Skill support aspect ratio control and output sizing for product mockups?

Yes, the Skill supports aspect ratio control and output sizing for product mockups. You can specify resolution and aspect ratio parameters within your text prompt to ensure the generated visual matches your specific formatting requirements.

Why use the Gemini API for style transfer instead of other image generation tools?

Using the Gemini API for style transfer allows you to apply conversational instructions to existing images directly within your workflow. Unlike standalone tools, it integrates multi-image composition and iterative refinement to shift styles without manual file exports.