image-generation

Generate and edit images from text prompts using Gemini and DALL-E models.

34|7|Updated Nov 29, 2025
One-click install
npx skills add https://github.com/jkitchin/skillz --skill image-generation-jkitchin
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: image-generation
Source: https://github.com/jkitchin/skillz/tree/main/skills/creative/image-generation
Command: npx skills add https://github.com/jkitchin/skillz --skill image-generation-jkitchin

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Enables generation, editing, and styling of images from natural language prompts, with multi-model support and iterative refinement.

Core Features & Use Cases

  • Text-to-Image: generate images from prompts with Gemini or DALL-E.
  • Image Editing: modify existing images via natural language instructions.
  • Style & Output: apply style transfers, create logos, product mockups, and variations.
  • Multi-Model Support: Gemini, DALL-E (3 and 2) with ground-truth references.

Quick Start

"Create a 4K product photo of a wireless earbud on a white background" or "Edit this image to place the object in a sunset scene."

Frequently Asked Questions about image-generation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate images from text prompts using AI?

Text-to-image generation converts natural language descriptions into visual assets using AI models. This Skill supports Gemini and DALL-E to create photorealistic photography, illustrations, logos, and product mockups from your prompts at configurable resolutions.

Can I edit existing images with text instructions?

Image editing via natural language lets you modify existing images—repositioning objects, changing backgrounds, or applying style transfers—by describing the changes you want. This Skill accepts image inputs and masking for precise, iterative refinement.

What AI models does this support for image generation?

This Skill supports Gemini 2.5 Flash, Gemini 3 Pro, DALL-E 3, and DALL-E 2, letting you choose the model that best fits your quality, speed, and cost requirements for each generation task.

How do I refine generated images across multiple turns?

Multi-turn refinement accepts reference images and iterative prompts, enabling rapid concepting and variation. You can adjust output resolution, aspect ratio, and style progressively until results match your vision.

Can I use this for professional contexts like product mockups and logos?

Yes. This Skill applies to business and professional use cases including product mockups, logo design, and rapid concepting. It supports batch generation at various resolutions for quick iteration across creative and commercial workflows.

What prompt-engineering practices should I follow for best results?

The Skill follows established prompt-engineering practices to maximize image quality. Provide clear, descriptive text prompts specifying style, composition, and details; use reference images to guide output; and iterate with refinements across multiple turns.