← Search

Sounak Mondal

7 accepted papers

2026

Personalized Image Descriptions from Attention Sequences

CVPR 2026

People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to substantial variability in image descriptions. However, existing models for personalized image description focus on lingu

Cited by 0SourcecodeScholar
2025

Few-shot Personalized Scanpath Prediction

CVPR 2025poster

A personalized model for scanpath prediction provides insights into the visual preferences and attention patterns of individual subjects. However, existing methods for training scanpath prediction models are data-intensive and cannot be effectively personalized to new individuals with only a few ava…

2025

Gaze-Language Alignment for Zero-Shot Prediction of Visual Search Targets from Human Gaze Scanpaths

ICCV 2025poster

Decoding human intent from eye gaze during a visual search task has become an increasingly important capability within augmented and virtual reality systems. However, gaze target prediction models used within such systems are constrained by the predefined target categories found within available gaz…

Cited by 0SourcePDFScholar
2024

Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following

ECCV 2024poster

"Training gaze following models requires a large number of images with gaze target coordinates annotated by human annotators, which is a laborious and inherently ambiguous process. We propose the first semi-supervised method for gaze following by introducing two novel priors to the task. We obtain t…

2024

Unifying Top-down and Bottom-up Scanpath Prediction Using Transformers

CVPR 2024poster

Most models of visual attention aim at predicting either top-down or bottom-up control as studied using different visual search and free-viewing tasks. In this paper we propose the Human Attention Transformer (HAT) a single model that predicts both forms of attention control. HAT uses a novel transf…

2023

Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human Attention

CVPR 2023poster

Predicting human gaze is important in Human-Computer Interaction (HCI). However, to practically serve HCI applications, gaze prediction models must be scalable, fast, and accurate in their spatial and temporal gaze predictions. Recent scanpath prediction models focus on goal-directed attention (sear…

2022

Target-Absent Human Attention

ECCV 2022poster

"The prediction of human gaze behavior is important for building human-computer interactive systems that can anticipate a user’s attention. Computer vision models have been developed to predict the fixations made by people as they search for target objects. But what about when the image has no targe…