CR-CLIP: Image-Text Contrastive Regression for Generalized Gaze Estimation
Yitong Zhu, Xurong Xie, Naiming Yao, Hui Chen, Feng Tian
Abstract
Gaze estimation methods typically encounter significant performance degradation in generalized tasks due to the domain mismatch between the source and target domains. Existing approaches attempt to utilize various domain generalization techniques. However, their generalization capabilities are limited since they are constrained to a single visual modality. Notably, large-scale contrastive language-image pre-training (CLIP) models have been widely applied to downstream visual tasks for their robust generalization capabilities, but the potential of CLIP for regression tasks has not been fully explored. To bridge this gap, we introduce a novel framework called CR-CLIP, which endows CLIP with the capability to generalize gaze estimation. Specifically, we convert gaze labels into textual descriptions and achieve alignment between images and text signals with gaze cues, thereby extracting generalized gaze-related features. To enhance the model’s understanding of the numerical relationships of gaze directions, we propose a novel regression loss function based on image-text similarity. Additionally, we fine-tune the model on the original gaze dataset, achieving high precision in generalized gaze estimation. Experimental results show that our proposed method achieves state-of-the-art performance on four generalized gaze estimation tasks.
BibTeX
@inproceedings{icassp2025_crclipimagetextc,
title = {CR-CLIP: Image-Text Contrastive Regression for Generalized Gaze Estimation},
author = {Yitong Zhu and Xurong Xie and Naiming Yao and Hui Chen and Feng Tian},
booktitle = {ICASSP 2025},
year = {2025}
}