2025
Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors
EMNLP 2025
Recent advancements in large video-language models have revolutionized video understanding tasks. However, their efficiency is significantly constrained by processing high volumes of visual tokens. Existing token compression strategies apply a fixed compression ratio, ignoring the variability in sem