2025
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
CVPR 2025poster
Multimodal language models (MLMs) still face challenges in fundamental visual perception tasks where specialized models excel. Tasks requiring reasoning about 3D structures benefit from depth estimation, and reasoning about 2D object instances benefits from object detection. Yet, MLMs can not produc…