Skip to main content
Live
Main content

Google ships Nano Banana 2 Lite and Gemini Omni Flash to developers

DeepMind's fastest image model generates in 4 seconds at $0.034 per 1K images; Omni Flash matches Veo 3.1 Fast at $0.10 per second of video.

Jaeden Schafer
Editor in Chief · · 5 min read
Google logo

Google DeepMind opened two new generative media models to developers today: Nano Banana 2 Lite, a text-to-image model that produces outputs in 4 seconds at $0.034 per 1K images, and Gemini Omni Flash, a video generation and editing model priced at $0.10 per second of output. Both ship through Google AI Studio, the Gemini API, and the Gemini Enterprise Agent Platform, with Nano Banana 2 Lite also landing inside AI Mode in Search, the Gemini app, NotebookLM, Google Photos, Stitch, Google Flow, and Google Ads on the same day.

The pricing is the point. At $0.034 per 1K images, Nano Banana 2 Lite is positioned for high-throughput pipelines where developers need to draft, iterate, or run low-bandwidth consumer features without burning budget on a heavier model. The 4-second latency is the other half of that pitch, aimed at interactive prototyping and any product where a user is waiting on the result.

We're making it easier to experiment and scale your ideas with Nano Banana 2 Lite, our fastest, most cost-efficient Gemini Image model, and Gemini Omni Flash for high-quality video generation and conversational editing.
Alisa Fortin, Product Manager, Google DeepMind

DeepMind is naming Nano Banana 2 Lite (gemini-3.1-flash-lite-image) the recommended replacement for the original Nano Banana, which shipped as gemini-2.5-flash-image. The company says the Lite tier holds onto prompt adherence, character consistency, and in-image text rendering despite the speed-cost optimization, though it sits below Nano Banana 2 (Gemini 3.1 Flash Image) and Nano Banana Pro (Gemini 3 Pro Image) on quality benchmarks. The family now spans four tiers, with Pro reserved for tasks where accuracy outweighs speed.

Key facts

  • 01Nano Banana 2 Lite generates text-to-image outputs in 4 seconds at $0.034 per 1K images, positioning it as DeepMind's speed-and-cost model.
  • 02Gemini Omni Flash is priced at $0.10 per second of video output, matching Veo 3.1 Fast, and supports 10-second generations at launch.
  • 03The Interactions API lets developers chain up to 3 sequential edits while maintaining session history across multi-turn workflows.
  • 04Nano Banana 2 Lite ships across Google AI Studio, the Gemini API, AI Mode in Search, the Gemini app, NotebookLM, Google Photos, Stitch, Flow, and Google Ads.
  • 05Omni Flash accepts video references up to 3 seconds, though DeepMind says they are not yet correctly processed by the model.

Gemini Omni Flash, first previewed at Google I/O, is the more ambitious of the two releases. It generates and edits video from a mix of text, image, and video inputs, and at $0.10 per second of output it matches Veo 3.1 Fast on price — meaning DeepMind is selling Omni's conversational editing and multimodal referencing without charging a premium over its existing fast-tier video model.

The selling features are conversational editing in natural language, multimodal referencing that mixes images, text, and video as control inputs, and text-to-action synchronization that ties on-screen graphics directly to video movement. DeepMind also leans on Gemini's broader knowledge base — history, biology, narrative logic — as a differentiator versus pure diffusion-based video systems.

Limitations are spelled out. Omni Flash currently caps video generation at 10 seconds, with longer durations promised later. Audio reference uploads and scene extension are not yet supported in the Gemini API for this model. Video references up to 3 seconds are accepted by the API schema but are not correctly processed at this time. And character consistency degrades during scene changes and panning movements, which DeepMind says it is working to improve.

The intended workflow chains the two models. A developer generates a still image with Nano Banana 2 Lite, then passes it to Omni Flash as a reference to animate into a video. The Interactions API maintains session history across that handoff and lets users stack up to three sequential edits in a single session, which is what turns the pair into something closer to an end-to-end creative tool rather than two disconnected endpoints.

DeepMind shipped three demo apps to anchor the pattern. Anywhere takes a selfie or uploaded photo and uses Nano Banana 2 Lite to place the subject at iconic landmarks, then animates the result with Omni Flash. Space Lift is an interior design demo that reimagines a room across aesthetics and produces a cinematic walkthrough. Omni product studio converts static product images into e-commerce video. All three are positioned as remixable starting points rather than finished products.

Related · from this week
Google brings Nano Banana and Veo to Google TV through a new Gemini tab
Jaeden Schafer · 4 min read →

Both models carry SynthID watermarking, and DeepMind is routing verification through the Gemini app, Gemini in Chrome, and Search. That's the company's standard provenance stack at this point, and it puts the burden of detection on viewers and platforms rather than on the generation step itself — a deliberate choice that lets the models ship without throughput penalties.

The open questions are the limitations DeepMind already named. Ten seconds is short for narrative video, character consistency across cuts is the hardest problem in the field and Omni is not yet solving it, and the 3-second video reference path is documented but broken. Product Manager Alisa Fortin said the team is working on character consistency, but until those fixes land, Omni's strongest use case is single-shot effects and short cinematic loops rather than continuous scenes.

Strategically, DeepMind is doing what OpenAI and Anthropic have not: pushing aggressive per-unit pricing on generative media at the same time it is fanning the models across every Google surface that touches a consumer. Nano Banana 2 Lite at $0.034 per 1K images undercuts most third-party image APIs by a wide margin, and chaining it into Omni Flash gives developers a vertically integrated image-to-video stack inside one vendor and one billing relationship. For startups building creative tools on top of someone else's models, that price-and-distribution combination is the real story — and the pressure it puts on standalone image and video generation companies is the part to watch over the next two quarters.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Google
Tools

Google brings Nano Banana and Veo to Google TV through a new Gemini tab

A Create button inside the Gemini tab lets viewers generate images and video clips by voice, starting on TCL sets in the US.

Jaeden Schafer4 min read
Google logo
Models

Google DeepMind's WeatherNext 3 delivers hourly forecasts at 5-kilometer resolution

The new model trains on live satellite data instead of physics simulations, cutting the six-hour lag that plagues traditional forecasting.

Jaeden Schafer5 min read
Google logo
Models

Google bakes computer use into Gemini 3.5 Flash as a native tool

DeepMind folds its standalone agent model into Flash, letting developers build agents that drive browsers, mobile apps and desktops via one API call.

Jaeden Schafer4 min read