Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding
Answering questions about long videos without processing every frame each time
A new system lets edge devices understand hours-long videos efficiently by separating what needs pictures from what doesn't. It runs one captioning pass upfront to create a searchable story skeleton, then only retrieves actual video frames when a question specifically needs visual details like text or appearance—cutting the computing work dramatically while keeping accuracy high.
Video analysis on phones, cameras, and other edge devices is bottlenecked by bandwidth and battery. This approach could enable real-time video understanding for surveillance, security, and accessibility apps without constantly streaming frames to the cloud or draining device power. The efficiency gains are concrete: you get strong accuracy with far less visual data flowing in and out.