Nano Banana 2 Lite & Gemini Omni Flash (Google)
Google’s launch post for two generative-media models in the gemini line — the first time this spoke has concrete IDs, latency, and per-unit pricing for the family’s image and video tiers (the May 2026 roundup only had marketing positioning, google-ai-updates-may-2026). T3 vendor blog; numbers are launch claims, treat as dated snapshots.
Nano Banana 2 Lite — text-to-image
- Model ID
gemini-3.1-flash-lite-image; the speed/cost tier of Google’s Nano Banana image line, and the replacement for the first-gen Nano Banana. - ~4-second generations at $0.034 per 1,000 images, pitched for “rapid ideation and high-velocity developer pipelines where speed and cost are the primary constraints.”
- Claims strong prompt adherence, character consistency, and legible in-image text — Google shows it on an Elo-vs-latency-vs-cost trade-off chart against competitors (the quality axis is an Elo score, same shape as the text llm-benchmarks).
Gemini Omni Flash — video generation & editing
- Model ID
gemini-omni-flash-preview; the cost-efficient tier of the Gemini Omni line, for video generation and conversational (natural-language) editing, with multimodal referencing (text + image- video inputs).
- $0.10 per second of video output — explicitly “the same as Veo 3.1 Fast,” so Google is pricing Omni Flash at parity with its existing fast video model. 10-second max generation length.
- Preview limits stated honestly: no audio references, no API scene-extension, video refs up to 3s accepted but not processed correctly, character consistency wobbles across scene changes.
The intended workflow
Google pitches the two together: generate a still with Nano Banana 2 Lite, then pass it as a reference to Omni Flash to animate it into video — an image→video pipeline, with the Interactions API supporting up to three sequential edits. Both ship on AI Studio, the Gemini API, and the Gemini Enterprise Agent Platform, rolling out to consumer surfaces (AI Mode in Search, the Gemini app, Google Flow).
Why it’s here, and a boundary note
These sit in gemini as the family’s generative-media tiers — the provider/model market this spoke owns. They also add a new pricing unit to the spoke ($/image, $/second-of-video) alongside the $/token text tiers in llm-api-pricing. Boundary watch: image/video generation is a distinct modality from the text/multimodal LLMs at this spoke’s core (and audio-gen was already carved out to speech-audio-wiki). For now it routes here as part of the Gemini family; if non-Google generative-media sources accumulate, a dedicated image/video-gen spoke is the spin-out candidate (see synthesis).
Related
gemini · google · google-ai-updates-may-2026 · llm-api-pricing · llm-benchmarks · synthesis