Models & Research

Google Gemini’s new agent-based video analysis cuts token usage by up to 88 percent

· September 2, 2026
Google Gemini’s new agent-based video analysis cuts token usage by up to 88 percent

What happened

Google has upgraded its Gemini video analysis models—Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite—with new agent-based video scanning. Instead of processing videos at a fixed frame rate, the models dynamically select which segments to analyze and adjust the resolution on the fly. This targeted approach reduces token usage by up to 88 percent while boosting accuracy, especially for videos that run multiple hours.

Why it matters

Video analysis on large-scale or long content typically demands vast computational resources. By cutting token consumption so sharply, Google lowers processing costs and speeds up workflows that handle long-distance surveillance, streaming archives, event logs, or other extended footage. The system’s ability to decide what to analyze means operators spend less on excessive frame-by-frame scanning, which can waste tokens on irrelevant or redundant data. This efficiency could reshape how companies deploy video AI for monitoring and content search, enabling cheaper, faster, and more reliable outcomes.

What to watch next

Observe how Google integrates this agent-based approach with existing video AI tools and if competitors follow suit to cut down on computational overhead. Also, be aware of how this method performs across diverse video types, like security footage, sports, or user-generated content. Accuracy improvements claim to be better with longer inputs, so verifying these gains in real-world deployments will be key. Finally, watch if token savings translate into lower cloud processing costs or new pricing models from video AI providers.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.