2025
ViPE: Visual Perception in Parameter Space for Efficient Video-Language Understanding
EMNLP 2025
Existing video-language models (Video-LLMs) typically rely on concatenating visual tokens with textual inputs for joint modeling. However, this token-level alignment leads to significant inefficiency, especially when scaling to long videos with dense visual inputs. In this work, we propose a video-t