Google released Gemini Omni 1.1 Flash, a production update to its native multimodal video model that puts more directorial control in developers’ hands. The model now reads up to 10 seconds of prior context when continuing a clip, up from a single final frame, and can extend scenes in 10-second increments to a cumulative 40 seconds.
Creators can pin a start and end frame and let the model generate the continuous shot between them, which powers camera orbits, dolly zooms and seamless loops. Up to three short video clips can be passed as references for character consistency. Google reports 360p previews render up to 60 percent faster at a third of the cost of 720p, making a draft-then-upscale loop the intended workflow, with finals upscaling to 4K.
Pricing runs $1.50 per million input tokens and $17.50 per million video output tokens, an effective $0.10 per second of 720p footage. Every generated video carries SynthID watermarking for provenance, and editing stays stateful through the Interactions API so changes preserve untouched elements of a prior clip.
Adobe Firefly, Figma Weave, GMI Cloud and Runway are already running the model in production, and it is live in Google Flow with scene extension in the Gemini app.