← Search

Haofei Wang

4 accepted papers

2026

ApET: Approximation-Error Guided Token Compression for Efficient VLMs

CVPR 2026

Recent Vision-Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibitive computational overhead and degrade inference efficiency. Prior studies typically relies on [CLS] attention or text-vision cross-attention to iden

Cited by 0SourcecodeScholar
2021

Generalizing Gaze Estimation With Outlier-Guided Collaborative Adaptation

ICCV 2021poster

Deep neural networks have significantly improved appearance-based gaze estimation accuracy. However, it still suffers from unsatisfactory performance when generalizing the trained model to new domains, e.g., unseen environments or persons. In this paper, we propose a plug-and-play gaze adaptation fr…

Cited by 71PDFcodeScholar