2026
FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging
ICLR 2026oral
Although Video Large Language Models (VLLMs) have shown remarkable capabilities in video understanding, they are required to process high volumes of visual tokens, causing significant computational inefficiency. Existing VLLMs acceleration frameworks usually compress spatial and temporal redundancy…