nano-banana-2

Generate Gemini 3.1 flash image previews from prompts via the inference.sh CLI.

4|1|Updated Feb 10, 2026
One-click install
npx skills add https://github.com/Sheshiyer/brandmint-oracle-aleph --skill nano-banana-2-sheshiyer
Or copy as Structured Prompt for Agent
Please help me install this Agent Skill.
Skill: nano-banana-2
Source: https://github.com/Sheshiyer/brandmint-oracle-aleph/tree/main/skills/external/inference-sh/upstream/ab546d072f1e/tools/image/nano-banana-2
Command: npx skills add https://github.com/Sheshiyer/brandmint-oracle-aleph --skill nano-banana-2-sheshiyer

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill enables rapid generation of Gemini 3.1 flash image previews from prompts using the inference.sh CLI, streamlining creative exploration.

Core Features & Use Cases

  • Text-to-image generation: produce Gemini 3.1 flash image previews from descriptive prompts.
  • Image editing and multi-input support: modify visuals and combine up to 14 input images for cohesive outputs.
  • Grounding via Google Search: optionally ground results with real-time context from the web to improve relevance.

Quick Start

Run the Nano Banana 2 skill to generate Gemini 3.1 flash image previews from a prompt.

Frequently Asked Questions about nano-banana-2

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I generate Gemini 3.1 flash image previews from a text prompt?

To generate Gemini 3.1 flash image previews from a text prompt, you can use a skill that runs the inference.sh CLI to process descriptive prompts and produce generated images with descriptive metadata.

Can I edit existing images and combine multiple inputs for text-to-image generation?

Yes, text-to-image generation workflows can edit existing visuals and combine up to 14 input images to create cohesive generated outputs using the inference.sh CLI.

Does image generation with Gemini 3.1 support web grounding via Google Search?

Yes, image generation with Gemini 3.1 supports optional web grounding via Google Search to provide real-time context and improve the relevance of generated image previews.

What is the best way to ground text-to-image outputs with real-time web context?

The best way to ground text-to-image outputs is by enabling web grounding via Google Search during generation, which applies real-time context to the image previews created through the inference.sh CLI.

Do I need the inference.sh CLI to run multi-image editing and generation?

Yes, you need the inference.sh CLI installed to run multi-image editing and text-to-image generation, as it processes the input prompts and up to 14 images to produce the final visual outputs.