# Google launches agentic video understanding in Gemini API

Models can dynamically inspect specific frames, audio, and transcripts instead of fixed frame rates, reducing token consumption by up to 88%.

- Cuts overall analysis costs by up to 66% while increasing accuracy by up to 7% on standard benchmarks
- Available for video uploads and YouTube links via Google AI Studio, the Gemini API, and Gemini Enterprise Agent Platform
- Supports Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite at standard token rates with no feature fee
- Enabled in API requests by setting the processing parameter to agentic
- Enables sub-second moment retrieval, high-FPS anomaly window resampling, and object counting in long videos

## Why it matters

Developers querying long-form video can extract context from multi-hour footage with higher accuracy while using fewer tokens and paying lower API costs.

## Sources

- [Google AI Studio: Google launches agentic video understanding in Gemini API](https://x.com/GoogleAIStudio/status/2094841307935957304)
- [Google DeepMind: Introducing agentic video understanding with Gemini](https://deepmind.google/blog/introducing-agentic-video-in-gemini)

---

Summarized by dstilled on 2026-09-01. https://dstilled.ai/story/69818eca-2ea7-4a81-a6d3-756624e32f67
