← Search

Yingying Ao

2 accepted papers

2025

Glance2Gaze: Efficient Vision-Language Models from Glance Fusion to Gaze Compression

NeurIPS 2025poster

Vision-language models heavily rely on visual representations, yet ensuring its efficiency remains a critical challenge. Most existing approaches focus on reducing visual tokens either at the visual encoder phase or during the LLM decoder stage. Inspired by human visual cognition, where an initial g…

Cited by 0SourceScholar
2024

CustomListener: Text-guided Responsive Interaction for User-friendly Listening Head Generation

CVPR 2024poster

Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion. The applications of listener agent generation in virtual interaction have promoted many works achieving diverse and fine-grained…

Cited by 4SourcePDFScholar