Google launches agentic video understanding in Gemini API
Models can dynamically inspect specific frames, audio, and transcripts instead of fixed frame rates, reducing token consumption by up to 88%.
- Cuts overall analysis costs by up to 66% while increasing accuracy by up to 7% on standard benchmarks
- Available for video uploads and YouTube links via Google AI Studio, the Gemini API, and Gemini Enterprise Agent Platform
- Supports Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite at standard token rates with no feature fee
- Enabled in API requests by setting the processing parameter to agentic
- Enables sub-second moment retrieval, high-FPS anomaly window resampling, and object counting in long videos
Developers querying long-form video can extract context from multi-hour footage with higher accuracy while using fewer tokens and paying lower API costs.

Sources
Read this as text
Back to the AI news