Google DeepMind released Gemini Omni 1.1 Flash today, a production-ready update to its generative video model that adds 4K upscaling, scene extension up to 40 seconds, and first-and-last-frame interpolation. The model is live through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, and Adobe, Figma, and GMI Cloud are already shipping it inside production tools. The update is aimed squarely at developers building video workflows, not at consumer prompt-and-pray toys.
The headline additions are control features that generative video has been missing. Scene extension now works in 10-second increments up to a cumulative 40 seconds, with the model analyzing 10 seconds of prior context to keep characters, lighting, and camera motion consistent. That is a step change from earlier Omni releases, which referenced only the final second of a clip before generating what came next.
“Omni now delivers studio-quality video production, including the ability to extend a scene, first and last frame interpolation, crisp 4K upscaling, faster prototyping, and more.”— Anish Nangia and Alisa Fortin, Product Managers, Google DeepMind
First and last frame conditioning is the other structural upgrade. Developers can now specify a starting frame and an ending frame, and Omni 1.1 Flash generates the continuous motion between them. Google DeepMind is pitching the feature at complex camera work — orbits, dolly zooms, seamless loops — where randomness in the middle of a shot has been the main reason AI video still looks like AI video.
Key facts
- 01Gemini Omni 1.1 Flash extends generated scenes up to a cumulative 40 seconds, in 10-second increments.
- 02The model can now reason over 10 seconds of prior context, up from the final second in earlier versions.
- 03360p drafts generate up to 60% faster and cost a third of standard 720p output.
- 04Final output can be upscaled to 1080p or 4K, with video references up to three seconds long.
- 05Adobe Firefly, Figma Weave, and GMI Cloud are already shipping the model in production.
The economics matter as much as the controls. Omni 1.1 Flash can generate 360p previews up to 60% faster than 720p output and at roughly a third of the cost, which turns iteration from an expensive gamble into cheap scouting. Once a draft is approved, the same model upscales to 1080p or 4K for finishing. That two-tier pattern — cheap drafts, expensive finals — has been standard in film pipelines for decades, and Google DeepMind is finally bringing it to generative video.
Multimodal input has also expanded. Developers can now pass up to three seconds of reference video alongside text prompts and still images, giving the model concrete visual anchors for character consistency and motion style. A demo in the announcement swaps three animated characters — a dog, an octopus, and a bear — into distinct reference dances filmed by human performers, keeping the choreography intact across a single continuous shot.
The API surface is deliberately thin. A scene-extension call takes a previous_interaction_id, a text prompt, and a resolution field, and the model handles context stitching internally. That design choice pushes complexity off the developer and onto the model, which is the right bet if Google DeepMind wants adoption inside creative software rather than inside research notebooks.
Adobe has integrated the model into Firefly for video editing, and Figma Weave is using it as one of the core generators on its canvas. GMI Cloud is running it through the Agent Platform API for production customers. Product managers Anish Nangia and Alisa Fortin, who authored the announcement, framed the release as the point at which Omni crosses from demo to deployment.
The competitive frame is unavoidable. Generative video has been the most contested frontier in AI over the past year, with OpenAI's Sora, Runway, Kling, and Meta's Movie Gen all pushing on quality and duration. Google DeepMind is not trying to win on wow-factor single clips; it is trying to win on the parts that matter for actual production — length, control, resolution, and cost per iteration. A 40-second cumulative ceiling is still short compared to any real narrative work, and the 360p draft tier is a tacit admission that 4K generation remains expensive enough that developers need a cheaper mode to explore.
The rougher edges are the ones this class of model still shares. Character identity can drift across a 40-second extension chain, physics in complex motion remains inconsistent, and 4K upscaling amplifies artifacts as often as it hides them. Google DeepMind has not published benchmark scores against competing video models, and the demos in the announcement are curated rather than randomly sampled. Developers will find the failure modes quickly once real workloads hit the API.
The interesting shift here is where video generation is heading as a product category. Text-to-video as a novelty is finished; the next race is over whether generative models can slot into existing creative pipelines with the controls that professional users demand. Adobe's Firefly integration and Figma Weave's canvas approach both suggest the same thesis — the model is a component, not the product. If Google DeepMind can hold the cost curve on drafts while pushing scene length and control features, Omni becomes the default backend for AI video inside the software people already use. That is a more valuable position than winning any single benchmark clip.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




