2026
MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
ICLR 2026poster
Despite advancements in Multimodal Large Language Models (MLLMs), their proficiency in fine-grained video motion understanding remains critically limited. They often lack inter-frame differencing and tend to average or ignore subtle visual cues. Furthermore, while visual prompting has shown potentia…