Google launched agentic video analysis in the Gemini API on September 1, 2026. The model can navigate the timeline and select frames, audio or transcripts instead of always using a fixed sampling sequence.
The announcement covered Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite; current documentation also lists 3.8 Flash. Google reports up to 88% lower token consumption in its tests. This is a task-dependent maximum, not guaranteed savings.
We see potential in locating a moment within a long lecture or maintenance recording. A person could first receive a suggested time segment, then return to the original video to verify it.
Evaluation needs to include missed events and response time. The documentation acknowledges that latency-sensitive short clips may be better suited to static processing.
Optimistic editorial estimate: a pilot using an organisation’s own archive could be assessed within 1–2 months. Reliability and ongoing costs cannot be predicted without testing.

Be the first to open the discussion.