gemini-imagegen

Generate and edit images from text prompts via the Gemini API.

3|Updated Feb 5, 2026
One-click install
npx skills add https://github.com/roach88/compound-engineering --skill gemini-imagegen-roach88
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: gemini-imagegen
Source: https://github.com/roach88/compound-engineering/tree/main/skills/gemini-imagegen
Command: npx skills add https://github.com/roach88/compound-engineering --skill gemini-imagegen-roach88

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires google-genai, Pillow, and includes scripts (resource) components.

What problem does it solve?

This tool enables automated image generation and editing through the Gemini API, eliminating manual, iterative design cycles and enabling rapid visuals from simple prompts.

Core Features & Use Cases

  • Text-to-image generation with configurable prompts, aspect ratios, and resolutions.
  • Image editing via natural-language instructions to modify existing visuals.
  • Image composition from multiple references to produce cohesive scenes or logos.
  • Interactive refinement through multi-turn prompts in a chat-like workflow.

Quick Start

Generate a 1:1 logo concept for Acme Corp using Gemini imagegen.

Frequently Asked Questions about gemini-imagegen

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate images from text prompts using the Gemini API?

This tool automates text-to-image generation by processing configurable text prompts, aspect ratios, and resolutions via the Gemini API to rapidly produce targeted visuals.

Can I edit existing images with natural language instructions in Gemini?

You can edit existing images in Gemini by providing natural-language instructions to programmatically modify and refine visuals without manual design iterations.

What is the best way to compose multiple image references into a single scene?

To compose multiple image references into a single scene, use the Gemini API to programmatically merge inputs and produce cohesive visuals or logos without manual editing.

Does this image generation method support multi-turn refinement workflows?

This approach supports multi-turn refinement workflows by using an interactive, chat-like process to iteratively adjust prompts and refine generated visuals.

Are configurable aspect ratios and resolutions available for programmatic image generation?

Configurable aspect ratios and resolutions are available for programmatic image generation, allowing exact output dimension specifications directly via the Gemini API.