← Search

Gregory Zelinsky

7 accepted papers

2026

Personalized Image Descriptions from Attention Sequences

CVPR 2026

People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to substantial variability in image descriptions. However, existing models for personalized image description focus on lingu

Cited by 0SourcecodeScholar
2024

Look Hear: Gaze Prediction for Speech-directed Human Attention

ECCV 2024poster

"For computer systems to effectively interact with humans using spoken language, they need to understand how the words being generated affect the users’ moment-by-moment attention. Our study focuses on the incremental prediction of attention as a person is seeing an image and hearing a referring exp…

2024

Unifying Top-down and Bottom-up Scanpath Prediction Using Transformers

CVPR 2024poster

Most models of visual attention aim at predicting either top-down or bottom-up control as studied using different visual search and free-viewing tasks. In this paper we propose the Human Attention Transformer (HAT) a single model that predicts both forms of attention control. HAT uses a novel transf…

2023

Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human Attention

CVPR 2023poster

Predicting human gaze is important in Human-Computer Interaction (HCI). However, to practically serve HCI applications, gaze prediction models must be scalable, fast, and accurate in their spatial and temporal gaze predictions. Recent scanpath prediction models focus on goal-directed attention (sear…

2022

Target-Absent Human Attention

ECCV 2022poster

"The prediction of human gaze behavior is important for building human-computer interactive systems that can anticipate a user’s attention. Computer vision models have been developed to predict the fixations made by people as they search for target objects. But what about when the image has no targe…

2020

Predicting Goal-Directed Human Attention Using Inverse Reinforcement Learning

CVPR 2020oral

Human gaze behavior prediction is important for behavioral vision and for computer vision applications. Most models mainly focus on predicting free-viewing behavior using saliency maps, but do not generalize to goal-directed behavior, such as when a person searches for a visual target object. We pro…

Cited by 136PDFcodeScholar
2015

Efficient Video Segmentation Using Parametric Graph Partitioning

ICCV 2015poster

Video segmentation is the task of grouping similar pixels in the spatio-temporal domain, and has become an important preprocessing step for subsequent video analysis. Most video segmentation and supervoxel methods output a hierarchy of segmentations, but while this provides useful multiscale informa…

Cited by 35PDFScholar