SURGE: Surprise-Guided Token Reduction for Efficient Video Understanding with VLMs
Videos contain rich information but also high redundancy, as consecutive frames often share similar backgrounds and predictable motions. Current video-language models (VLMs) are unable to exploit this redundancy and therefore perform a significant amount of superfluous computation, processing thousa…