MODEL SIGNAL
Nano Banana 2 Lite
Google introduces Gemini 3.1 Flash-Lite Image, pushing text-to-image generation into the four-second latency tier.
Bottom line
Released by Google on June 30, 2026, Nano Banana 2 Lite—officially addressable via API as gemini-3.1-flash-lite-image—is optimized for high-throughput text-to-image generation and editing. The model outputs 1K-resolution images in approximately four seconds, establishing a faster, cost-efficient tier within the Gemini image family.
Signal
The primary signal lies in the unit economics and generation speed. According to Google's announcement, the model achieves a text-to-image generation latency of roughly four seconds for a standard 1K-resolution output (approximately 1024px, or ~1MP).
On the pricing front, the Gemini API documentation indicates input costs of $0.25 per 1M tokens and output costs of $1.50 per 1M tokens. Because each 1K-resolution image consumes 1,120 output tokens, the effective standard API cost translates to about $0.0336–$0.034 per single generated image. The model is actively rolling out across multiple Google surfaces, including Google AI Studio, the Gemini API, and the Gemini Enterprise Agent Platform, as well as consumer touchpoints like the Gemini app and AI Mode in Search.
Noise
The dual branding between consumer and developer surfaces may cause initial routing confusion. While the model is publicly discussed under the "Nano Banana 2 Lite" moniker, the official architectural name and API model ID is gemini-3.1-flash-lite-image. Operators must ensure they are pointing to the correct 3.1 Flash-Lite endpoint to realize the stated latency and cost profiles.
Where it fits
Nano Banana 2 Lite is designed for high-volume, low-latency image generation workflows. The operator read is that by cutting generation time to roughly four seconds, this model effectively shifts text-to-image capabilities from asynchronous, background-task queues into synchronous, near-real-time user flows.
The directional implication for developers is a widened aperture for interactive media applications. If the provider's latency and cost facts hold at production scale, this architecture equips teams to embed dynamic image generation directly into chat interfaces, rapid prototyping tools, and high-throughput content pipelines without incurring the heavy UX penalty of standard diffusion wait times.