Research
Google Adds Agentic Video Analysis to Gemini With Up to 88% Fewer Tokens
Gemini can now decide which moments, frame rates, audio and transcripts to inspect rather than processing video at a fixed sampling rate. Google says the approach cuts token use by up to 88%, cost by up to 66% and improves accuracy by up to 7%.
By Michael G ·

Google has introduced agentic video understanding for Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite, allowing the models to decide which moments and modalities to inspect instead of ingesting every video at a fixed sampling rate. Google says the method can reduce token consumption by up to 88%, lower analysis cost by up to 66% and improve accuracy by up to 7%. The change treats video analysis as an active search problem rather than a large file to summarize.
Conventional multimodal pipelines sample video at a fixed number of frames per second, often one. That keeps computation predictable but misses brief actions and wastes tokens on uneventful sections. An hour-long lecture may contain one answer near the end. A safety camera may show a critical movement lasting less than a second. One sampling policy is poorly matched to both tasks.
Gemini's new loop can search transcripts, audio and visual frames, then load a relevant segment at a different speed or frame rate. The model may scan broadly, identify a candidate interval and revisit it in greater detail. Developers could build that orchestration themselves, but moving the selection loop into the model reduces application code and allows the system to adapt its strategy to each question.
The feature is available for uploaded files and YouTube videos through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. It uses standard token pricing without a separate feature fee. Developers enable it by setting video processing to agentic. Google plans to bring the capability to the Gemini app and eventually use it in Ask YouTube.
Dynamic Sampling Changes the Error Pattern
Fixed sampling fails visibly when an important event happens between frames. Agentic sampling can recover by increasing temporal resolution around suspected motion. It creates a different failure: the model may never inspect the right segment because its first search chose poorly. Efficiency depends on selecting evidence, and selection can encode bias. Evaluation should measure missed moments as well as the accuracy of answers based on moments the model found.

Google highlights sub-second retrieval, long-form search, anomaly detection and counting. Each demands a different observation strategy. Precise editing needs cut boundaries. A lecture question may rely on transcript search and one visual demonstration. Industrial anomaly detection may require comparing motion across repeated cycles. A single agent loop can coordinate those tools, but the output should expose which intervals were reviewed so users can audit the path.
The reported cost reduction is especially important for multi-hour media. Static processing can consume millions of tokens or force developers to compress the video before analysis. If agentic selection preserves accuracy while touching a fraction of the content, archives that were economically impractical become searchable. Broadcasters, manufacturers, educators and security teams could analyze far more footage without building custom retrieval systems for every collection.
Benchmarks offer a controlled comparison, but real footage adds camera cuts, poor audio, overlays, multiple languages and uncertain timestamps. The model may use a clean transcript as a shortcut in tests and struggle when speech is noisy or the answer is purely visual. Customers should test representative recordings and include cases where the transcript contradicts the frame. Multimodal understanding matters precisely when one channel is incomplete.
Lower Cost Expands the Governance Surface
Cheaper analysis will make it easier to process meetings, classrooms, workplaces and public spaces. That increases privacy and labor concerns. A system capable of finding one gesture across thousands of hours can support accessibility or quality control. It can also make pervasive monitoring affordable. Organizations need retention limits, purpose restrictions and access logs before pointing agentic search at footage collected for another reason.

Video answers should include temporal citations. A text response that says an event occurred is weaker than a link to the relevant interval and the frames the model used. Citations also reveal when the system repeatedly samples the same misleading shot or ignores context just outside the selected window. Google can make agentic analysis more trustworthy by returning the retrieval trace as a first-class output rather than an internal implementation detail.
Developers must budget for variable computation. Static sampling has predictable cost based on duration and frame rate. An agent may revisit difficult sections several times. Average cost can fall while the tail for hard questions remains high. Production systems need limits on tool calls, elapsed time and segment retrieval, along with a graceful result when the search budget ends before the model is confident.
The same variability affects latency. A simple transcript question may finish quickly, while counting rapid repeated movements requires several passes. Applications should not promise one response time for every query. They can offer a fast mode for broad search and a deliberate mode for evidence-intensive review, exposing the tradeoff rather than hiding it behind one button.
Video Models Are Becoming Media Operators
The system is no longer only recognizing content. It is deciding how to watch. That resembles the work of an editor, analyst or investigator who scrubs a timeline, listens for a phrase and slows a moment down. The analogy has limits because the model lacks human context and accountability. It is still useful for architecture: high-quality video analysis requires a plan for gathering evidence, not only a larger context window.
Google's rollout to YouTube could bring the capability to billions of users and an enormous structured video corpus. It may improve questions about demonstrations, lectures and product reviews. Creators will want to know how their footage is sampled and whether answers direct viewers back to the source. A system that extracts the useful moment without preserving attribution could reduce traffic even as it improves discovery.
Caching could extend the savings for repeated questions. An enterprise may ask many users about the same lecture, inspection recording or meeting. The platform should clarify whether segment indexes and intermediate representations are reused, how long they persist and whether a new upload produces a new cache. Reuse can lower cost while creating another stored derivative of potentially sensitive media.
Live video is an obvious next step and a harder one. An agent deciding what to inspect cannot rewind an event that was never sampled or retained. Real-time systems need a rolling buffer, event triggers and a policy for how much footage remains available. Latency and privacy requirements may favor local preprocessing even when the final reasoning occurs in the cloud.
Evaluation should include adversarial editing. A video can hide instructions in one frame, place misleading captions over footage or splice events from different times. An agent that actively searches may be more likely to find a hidden signal and more likely to follow it. Tool-use policies should treat media as untrusted input and prevent a frame or transcript from silently changing the system's objective.
Accessibility applications may benefit from efficient long-video search, including navigation within recorded lectures or finding visual demonstrations described poorly in transcripts. Those products should expose controls that work with screen readers and keyboard navigation. A more capable retrieval engine does not create access if the interface for reviewing the result remains difficult to use.
Agentic video understanding is a practical advance because it attacks the real bottleneck: most frames are irrelevant to a particular question. Its success will depend on whether the model can prove that it watched the right evidence, not merely whether it uses fewer tokens. The strongest result is a lower bill, a better answer and a clear path back to the seconds that justify it.
Topics: Google, Gemini, video understanding, multimodal AI, token efficiency