← Search

Minh Hoai

43 accepted papers

2026

Personalized Image Descriptions from Attention Sequences

CVPR 2026

People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to substantial variability in image descriptions. However, existing models for personalized image description focus on lingu

Cited by 0SourcecodeScholar
2025

Few-shot Personalized Scanpath Prediction

CVPR 2025poster

A personalized model for scanpath prediction provides insights into the visual preferences and attention patterns of individual subjects. However, existing methods for training scanpath prediction models are data-intensive and cannot be effectively personalized to new individuals with only a few ava…

2025

Multi-view Gaze Target Estimation

ICCV 2025poster

This paper presents a method that utilizes multiple camera views for the gaze target estimation (GTE) task. The approach integrates information from different camera views to improve accuracy and expand applicability, addressing limitations in existing single-view methods that face challenges such a…

Cited by 0SourcePDFScholar
2025

Region-Level Data Attribution for Text-to-Image Generative Models

ICCV 2025poster

Data attribution in text-to-image generative models is a crucial yet underexplored problem, particularly at the regional level, where identifying the most influential training regions for generated content can enhance transparency, copyright protection, and error diagnosis. Existing data attribution…

2024

Blur2Blur: Blur Conversion for Unsupervised Image Deblurring on Unknown Domains

CVPR 2024poster

This paper presents an innovative framework designed to train an image deblurring algorithm tailored to a specific camera device. This algorithm works by transforming a blurry input image which is challenging to deblur into another blurry image that is more amenable to deblurring. The transformation…

2024

Count What You Want: Exemplar Identification and Few-Shot Counting of Human Actions in the Wild

AAAI 2024technical

This paper addresses the task of counting human actions of interest using sensor data from wearable devices. We propose a novel exemplar-based framework, allowing users to provide exemplars of the actions they want to count by vocalizing predefined sounds ``one'', ``two'', and ``three''. Our method…

2024

Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following

ECCV 2024poster

"Training gaze following models requires a large number of images with gaze target coordinates annotated by human annotators, which is a laborious and inherently ambiguous process. We propose the first semi-supervised method for gaze following by introducing two novel priors to the task. We obtain t…

2024

Error Detection in Egocentric Procedural Task Videos

CVPR 2024poster

We present a new egocentric procedural error dataset containing videos with various types of errors as well as normal videos and propose a new framework for procedural error detection using error-free training videos only. Our framework consists of an action segmentation model and a contrastive step…

Cited by 15SourcePDFScholar
2024

HOIST-Former: Hand-held Objects Identification Segmentation and Tracking in the Wild

CVPR 2024poster

We address the challenging task of identifying segmenting and tracking hand-held objects which is crucial for applications such as human action segmentation and performance evaluation. This task is particularly challenging due to heavy occlusion rapid motion and the transitory nature of objects bein…

Cited by 3SourcePDFScholar
2024

HanDiffuser: Text-to-Image Generation With Realistic Hand Appearances

CVPR 2024poster

Text-to-image generative models can generate high-quality humans but realism is lost when generating hands. Common artifacts include irregular hand poses shapes incorrect numbers of fingers and physically implausible finger orientations. To generate images with realistic hands we propose a novel dif…

Cited by 27SourcePDFScholar
2024

Look Hear: Gaze Prediction for Speech-directed Human Attention

ECCV 2024poster

"For computer systems to effectively interact with humans using spoken language, they need to understand how the words being generated affect the users’ moment-by-moment attention. Our study focuses on the incremental prediction of attention as a person is seeing an image and hearing a referring exp…

2024

Unifying Top-down and Bottom-up Scanpath Prediction Using Transformers

CVPR 2024poster

Most models of visual attention aim at predicting either top-down or bottom-up control as studied using different visual search and free-viewing tasks. In this paper we propose the Human Attention Transformer (HAT) a single model that predicts both forms of attention control. HAT uses a novel transf…

2023

Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human Attention

CVPR 2023poster

Predicting human gaze is important in Human-Computer Interaction (HCI). However, to practically serve HCI applications, gaze prediction models must be scalable, fast, and accurate in their spatial and temporal gaze predictions. Recent scanpath prediction models focus on goal-directed attention (sear…

2023

HyperCUT: Video Sequence From a Single Blurry Image Using Unsupervised Ordering

CVPR 2023poster

We consider the challenging task of training models for image-to-video deblurring, which aims to recover a sequence of sharp images corresponding to a given blurry image input. A critical issue disturbing the training of an image-to-video model is the ambiguity of the frame ordering since both the f…

2022

Forward Propagation, Backward Regression, and Pose Association for Hand Tracking in the Wild

CVPR 2022poster

We propose HandLer, a novel convolutional architecture that can jointly detect and track hands online in unconstrained videos. HandLer is based on Cascade-RCNNwith additional three novel stages. The first stage is Forward Propagation, where the features from frame t-1 are propagated to frame t based…

Cited by 13PDFcodeScholar
2022

Target-Absent Human Attention

ECCV 2022poster

"The prediction of human gaze behavior is important for building human-computer interactive systems that can anticipate a user’s attention. Computer vision models have been developed to predict the fixations made by people as they search for target objects. But what about when the image has no targe…

2022

Whose Hands Are These? Hand Detection and Hand-Body Association in the Wild

CVPR 2022poster

We study a new problem of detecting hands and finding the location of the corresponding person for each detected hand. This task is helpful for many downstream tasks such as hand tracking and hand contact estimation. Associating hands with people is challenging in unconstrained conditions since mult…

Cited by 26PDFcodeScholar
2021

Dictionary-Guided Scene Text Recognition

CVPR 2021poster

Language prior plays an important role in the way humans perceive and recognize text in the wild. In this work, we present an approach to train and use scene text recognition models by exploiting multiple clues from a language reference. Current scene text recognition methods have used lexicons to i…

Cited by 74PDFcodeScholar
2021

Lipstick Ain't Enough: Beyond Color Matching for In-the-Wild Makeup Transfer

CVPR 2021poster

Makeup transfer is the task of applying on a source face the makeup style from a reference image. Real-life makeups are diverse and wild, which cover not only color-changing but also patterns, such as stickers, blushes, and jewelries. However, existing works overlooked the latter components and conf…

Cited by 77PDFcodeScholar
2021

Localization in the Crowd with Topological Constraints

AAAI 2021technical

We address the problem of crowd localization, i.e., the prediction of dots corresponding to people in a crowded scene. Due to various challenges, a localization method is prone to spatial semantic errors, i.e., predicting multiple dots within a same person or collapsing multiple dots in a cluttered…

2021

Toward Realistic Single-View 3D Object Reconstruction With Unsupervised Learning From Multiple Images

ICCV 2021poster

Recovering the 3D structure of an object from a single image is a challenging task due to its ill-posed nature. One approach is to utilize the plentiful photos of the same object category to learn a strong 3D shape prior for the object. This approach has successfully been demonstrated by a recent wo…

Cited by 11PDFcodeScholar
2020

Learning Visual Emotion Representations From Web Data

CVPR 2020poster

We present a scalable approach for learning powerful visual features for emotion recognition. A critical bottleneck in emotion recognition is the lack of large scale datasets that can be used for learning visual emotion features. To this end, we curate a webly derived large scale dataset, StockEmoti…

Cited by 50PDFScholar
2020

Predicting Goal-Directed Human Attention Using Inverse Reinforcement Learning

CVPR 2020oral

Human gaze behavior prediction is important for behavioral vision and for computer vision applications. Most models mainly focus on predicting free-viewing behavior using saliency maps, but do not generalize to goal-directed behavior, such as when a person searches for a visual target object. We pro…

Cited by 136PDFcodeScholar
2019

Contextual Attention for Hand Detection in the Wild

ICCV 2019poster

We present Hand-CNN, a novel convolutional network architecture for detecting hand masks and predicting hand orientations in unconstrained images. Hand-CNN extends MaskRCNN with a novel attention mechanism to incorporate contextual cues in the detection process. This attention mechanism can be imple…

Cited by 77PDFcodeScholar
2019

GIF2Video: Color Dequantization and Temporal Interpolation of GIF Images

CVPR 2019poster

Graphics Interchange Format (GIF) is a highly portable graphics format that is ubiquitous on the Internet. Despite their small sizes, GIF images often contain undesirable visual artifacts such as flat color regions, false contours, color shift, and dotted patterns. In this paper, we propose GIF2Vide…

Cited by 30PDFScholar
2018

A+D Net: Training a Shadow Detector with Adversarial Shadow Attenuation

ECCV 2018poster

We propose a novel GAN-based framework for detecting shadows in images, in which a shadow detection network (D-Net) is trained together with a shadow attenuation network (A-Net) that generates adversarial training examples. The A-Net modifies the original training images constrained by a simplified…

Cited by 142SourcePDFScholar
2018

Good View Hunting: Learning Photo Composition From Dense View Pairs

CVPR 2018poster

Finding views with good photo composition is a challenging task for machine learning methods. A key difficulty is the lack of well annotated large scale datasets. Most existing datasets only provide a limited number of annotations for good views, while ignoring the comparative nature of view select…

Cited by 110SourcePDFScholar
2017

Shadow Detection With Conditional Generative Adversarial Networks

ICCV 2017oral

We introduce scGAN, a novel extension of conditional Generative Adversarial Networks (GAN) tailored for the challenging problem of shadow detection in images. Previous methods for shadow detection focus on learning the local appearance of shadow regions, while using limited local context reasoning i…

Cited by 245PDFScholar