2025
Glance2Gaze: Efficient Vision-Language Models from Glance Fusion to Gaze Compression
NeurIPS 2025poster
Vision-language models heavily rely on visual representations, yet ensuring its efficiency remains a critical challenge. Most existing approaches focus on reducing visual tokens either at the visual encoder phase or during the LLM decoder stage. Inspired by human visual cognition, where an initial g…