What problem does it solve?
Provide a unified, production-ready interface to generate and edit images using multiple provider APIs, removing the overhead of handling provider-specific APIs, model selection, aspect ratios, reference-image edits, and environment configuration for everyday creative and automation tasks.
Core Features & Use Cases
- Multi-provider Support: Works with Google (Gemini/Imagen), OpenAI (GPT Image), DashScope (Alibaba), and Replicate with provider-specific handling.
- Reference Images & Edits: Support for multimodal reference-image editing when provider/model allows, plus automatic provider selection when references are supplied.
- Quality, Size & Aspect Control: Preset quality (normal/2k), explicit imageSize, and aspect ratio mapping with sensible defaults and fallbacks.
- Operational Safety & Preferences: Enforces a first-time EXTEND.md preference setup before generation, resolves models from CLI/EXTEND.md/env, displays chosen provider/model, retries transient failures, and offers optional parallel batch generation and R2 upload utilities.
- Use Case Examples: Create article covers at 2k, batch-generate social thumbnails, or perform reference-based edits using a Gemini multimodal model.
Quick Start
Generate a 2k landscape image of a golden retriever running on a beach and save it as out.png using the Google provider.