Google launches agentic video understanding in Gemini API

Models can dynamically inspect specific frames, audio, and transcripts instead of fixed frame rates, reducing token consumption by up to 88%.

Developers querying long-form video can extract context from multi-hour footage with higher accuracy while using fewer tokens and paying lower API costs.

Google launches agentic video understanding in Gemini API

Sources

Read this as text

Back to the AI news