Agents that sit through long videos on Gemini Flash no longer have to pay for every second up front. Google has started offering an agentic video understanding mode in which the model scouts a recording first and then reads only the stretches that answer the task, a break from the old habit of pulling an entire timeline into context at one frame per second.
Old behavior charged the same toll to every request. Feed the system a 90-minute lecture and the fixed single-pass pipeline carried the whole thing into context, whether the ask was a rough summary or a hunt for one pricing slide. Teams could pre-cut footage to control cost, but then they gambled that the interesting detail survived the cut.
In the new loop Gemini decides what to watch, how fast, and through which senses. Google says the switch removes up to 88 percent of the video tokens, cuts cost by as much as 66 percent, and nudges accuracy up 7 percent on standard video benchmarks.
The capability is a cloud feature rather than a download: no weights are published and there is nothing to self-host. Google routes it through the Gemini API in AI Studio and through the enterprise agent platform, where it accepts uploaded files and public YouTube links alike. Billing runs on ordinary Gemini token prices, with no add-on fee. Static one-frame-per-second handling stays as the default for teams that prefer predictable spend.
Developers building agents that digest meetings, lectures or camera feeds are the immediate winners. The same navigate-then-reason logic is a natural fit for the larger members of the Gemini line, so this is likely the first tier of a broader change.