Skip to main content
Live
Main content

Google DeepMind ships Gemini Omni 1.1 Flash with 4K video and scene extension

The Omni update pushes generative video to 4K, extends clips to 40 seconds, and cuts draft costs to a third at 360p.

Jaeden Schafer
Editor in Chief · · 5 min read
Google logo

Google DeepMind released Gemini Omni 1.1 Flash today, a production-ready update to its generative video model that adds 4K upscaling, scene extension up to 40 seconds, and first-and-last-frame interpolation. The model is live through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, and Adobe, Figma, and GMI Cloud are already shipping it inside production tools. The update is aimed squarely at developers building video workflows, not at consumer prompt-and-pray toys.

The headline additions are control features that generative video has been missing. Scene extension now works in 10-second increments up to a cumulative 40 seconds, with the model analyzing 10 seconds of prior context to keep characters, lighting, and camera motion consistent. That is a step change from earlier Omni releases, which referenced only the final second of a clip before generating what came next.

Omni now delivers studio-quality video production, including the ability to extend a scene, first and last frame interpolation, crisp 4K upscaling, faster prototyping, and more.
Anish Nangia and Alisa Fortin, Product Managers, Google DeepMind

First and last frame conditioning is the other structural upgrade. Developers can now specify a starting frame and an ending frame, and Omni 1.1 Flash generates the continuous motion between them. Google DeepMind is pitching the feature at complex camera work — orbits, dolly zooms, seamless loops — where randomness in the middle of a shot has been the main reason AI video still looks like AI video.

Key facts

  • 01Gemini Omni 1.1 Flash extends generated scenes up to a cumulative 40 seconds, in 10-second increments.
  • 02The model can now reason over 10 seconds of prior context, up from the final second in earlier versions.
  • 03360p drafts generate up to 60% faster and cost a third of standard 720p output.
  • 04Final output can be upscaled to 1080p or 4K, with video references up to three seconds long.
  • 05Adobe Firefly, Figma Weave, and GMI Cloud are already shipping the model in production.

The economics matter as much as the controls. Omni 1.1 Flash can generate 360p previews up to 60% faster than 720p output and at roughly a third of the cost, which turns iteration from an expensive gamble into cheap scouting. Once a draft is approved, the same model upscales to 1080p or 4K for finishing. That two-tier pattern — cheap drafts, expensive finals — has been standard in film pipelines for decades, and Google DeepMind is finally bringing it to generative video.

Multimodal input has also expanded. Developers can now pass up to three seconds of reference video alongside text prompts and still images, giving the model concrete visual anchors for character consistency and motion style. A demo in the announcement swaps three animated characters — a dog, an octopus, and a bear — into distinct reference dances filmed by human performers, keeping the choreography intact across a single continuous shot.

The API surface is deliberately thin. A scene-extension call takes a previous_interaction_id, a text prompt, and a resolution field, and the model handles context stitching internally. That design choice pushes complexity off the developer and onto the model, which is the right bet if Google DeepMind wants adoption inside creative software rather than inside research notebooks.

Adobe has integrated the model into Firefly for video editing, and Figma Weave is using it as one of the core generators on its canvas. GMI Cloud is running it through the Agent Platform API for production customers. Product managers Anish Nangia and Alisa Fortin, who authored the announcement, framed the release as the point at which Omni crosses from demo to deployment.

The competitive frame is unavoidable. Generative video has been the most contested frontier in AI over the past year, with OpenAI's Sora, Runway, Kling, and Meta's Movie Gen all pushing on quality and duration. Google DeepMind is not trying to win on wow-factor single clips; it is trying to win on the parts that matter for actual production — length, control, resolution, and cost per iteration. A 40-second cumulative ceiling is still short compared to any real narrative work, and the 360p draft tier is a tacit admission that 4K generation remains expensive enough that developers need a cheaper mode to explore.

Related · from this week
Google launches Gemini 3.5 Transcribe with 2.6% word error rate
Jaeden Schafer · 5 min read →

The rougher edges are the ones this class of model still shares. Character identity can drift across a 40-second extension chain, physics in complex motion remains inconsistent, and 4K upscaling amplifies artifacts as often as it hides them. Google DeepMind has not published benchmark scores against competing video models, and the demos in the announcement are curated rather than randomly sampled. Developers will find the failure modes quickly once real workloads hit the API.

The interesting shift here is where video generation is heading as a product category. Text-to-video as a novelty is finished; the next race is over whether generative models can slot into existing creative pipelines with the controls that professional users demand. Adobe's Firefly integration and Figma Weave's canvas approach both suggest the same thesis — the model is a component, not the product. If Google DeepMind can hold the cost curve on drafts while pushing scene length and control features, Omni becomes the default backend for AI video inside the software people already use. That is a more valuable position than winning any single benchmark clip.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Google logo
Models

Google launches Gemini 3.5 Transcribe with 2.6% word error rate

The new speech-to-text model hits a 4.0% streaming WER, cuts time to final transcription by 70%, and supports over 85 languages.

Jaeden Schafer5 min read
Inherent's Faraday agent beats Claude and GPT-5.5 at replicating research on a 27B model
Models

Inherent's Faraday agent beats Claude and GPT-5.5 at replicating research on a 27B model

The London lab, fresh off a $50M seed, says its DeepMind-alumni-built agent matches frontier systems using a fraction of the parameters.

Jaeden Schafer5 min read
Google logo
Models

Google bakes computer use into Gemini 3.5 Flash as a native tool

DeepMind folds its standalone agent model into Flash, letting developers build agents that drive browsers, mobile apps and desktops via one API call.

Jaeden Schafer4 min read