← Search

Haoxiang Li

27 accepted papers

2025

CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models

ICLR 2025poster

Virtual try-on methods based on diffusion models achieve realistic effects but often require additional encoding modules, a large number of training parameters, and complex preprocessing, which increases the burden on training and inference. In this work, we re-evaluate the necessity of additional m…

2025

GeoRemover: Removing Objects and Their Causal Visual Artifacts

NeurIPS 2025spotlight

Towards intelligent image editing, object removal should eliminate both the target object and its causal visual artifacts, such as shadows and reflections. However, existing image appearance-based methods either follow strictly mask-aligned training and fail to remove these casual effects which are…

Cited by 0SourcecodeScholar
2024

Enhancing Implicit Shape Generators Using Topological Regularizations

ICML 2024poster

A fundamental problem in learning 3D shapes generative models is that when the generative model is simply fitted to the training data, the resulting synthetic 3D models can present various artifacts. Many of these artifacts are topological in nature, e.g., broken legs, unrealistic thin structures, a…

Cited by 1SourcePDFScholar
2024

SAM-DEBLUR: Let Segment Anything Boost Image Deblurring

ICASSP 2024accepted

Image deblurring is a critical task in the field of image restoration, aiming to eliminate blurring artifacts. However, the challenge of addressing non-uniform blurring leads to an ill-posed problem, which limits the generalization performance of existing deblurring models. To solve the problem, we…

Cited by 0SourceScholar
2023

Flexible Visual Recognition by Evidential Modeling of Confusion and Ignorance

ICCV 2023poster

In real-world scenarios, typical visual recognition systems could fail under two major causes, i.e., the misclassification between known classes and the excusable misbehavior on unknown-class images. To tackle these deficiencies, flexible visual recognition should dynamically predict multiple classe…

Cited by 5PDFScholar
2023

Implicit Autoencoder for Point-Cloud Self-Supervised Representation Learning

ICCV 2023poster

This paper advocates the use of implicit surface representation in autoencoder-based self-supervised 3D representation learning. The most popular and accessible 3D representation, i.e., point clouds, involves discrete samples of the underlying continuous 3D surface. This discretization process intro…

Cited by 65PDFcodeScholar
2023

Weakly-Guided Self-Supervised Pretraining for Temporal Activity Detection

AAAI 2023technical

Temporal Activity Detection aims to predict activity classes per frame, in contrast to video-level predictions in Activity Classification (i.e., Activity Recognition). Due to the expensive frame-level annotations required for detection, the scale of detection datasets is limited. Thus, commonly, pre…

2022

Breadcrumbs: Adversarial Class-Balanced Sampling for Long-Tailed Recognition

ECCV 2022poster

"The problem of long-tailed recognition, where the number of examples per class is highly unbalanced, is considered. While training with class-balanced sampling has been shown effective for this problem, it is known to over-fit to few-shot classes. It is hypothesized that this is due to the repeated…

2021

GistNet: A Geometric Structure Transfer Network for Long-Tailed Recognition

ICCV 2021poster

The problem of long-tailed recognition, where the number of examples per class is highly unbalanced, is considered. It is hypothesized that the well known tendency of standard classifier training to overfit to popular classes can be exploited for effective transfer learning. Rather than eliminating…

Cited by 64PDFScholar
2021

Learning Dynamics via Graph Neural Networks for Human Pose Estimation and Tracking

CVPR 2021poster

Multi-person pose estimation and tracking serve as crucial steps for video understanding. Most state-of-the-art approaches rely on first estimating poses in each frame and only then implementing data association and refinement. Despite the promising results achieved, such a strategy is inevitably pr…

Cited by 98PDFScholar
2018

A Modulation Module for Multi-task Learning with Applications in Image Retrieval

ECCV 2018poster

Multi-task learning has been widely adopted in many computer vision tasks to improve overall computation efficiency or boost the performance of individual tasks, under the assumption that those tasks are correlated and complementary to each other. However, the relationships between the tasks are com…

2018

Active Object Perceiver: Recognition-Guided Policy Learning for Object Searching on Mobile Robots

IROS 2018poster

We study the problem of learning a navigation policy for a robot to actively search for an object of interest in an indoor environment solely from its visual inputs. While scene-driven visual navigation has been widely studied, prior efforts on learning navigation policies for robots to find objects…

Cited by 58SourceScholar
2018

Contemplating Visual Emotions: Understanding and Overcoming Dataset Bias

ECCV 2018poster

While machine learning approaches to visual emotion recognition offer great promise, current methods consider training and testing models on small scale datasets covering limited visual emotion concepts. Our analysis identifies an important but long overlooked issue of existing visual emotion benchm…

Cited by 106SourcePDFScholar
2018

Deep Face Detector Adaptation Without Negative Transfer or Catastrophic Forgetting

CVPR 2018poster

Arguably, no single face detector fits all real-life scenarios. It is often desirable to have some built-in schemes for a face detector to automatically adapt, e.g., to a particular user's photo album (the target domain). We propose a novel face detector adaptation approach that works as long as the…

Cited by 14SourcePDFScholar
2017

Learning Dense Facial Correspondences in Unconstrained Images

ICCV 2017poster

We present a minimalistic but effective neural network that computes dense facial correspondences in highly unconstrained RGB images. Our network learns a per-pixel flow and a matchability mask between 2D input photographs of a person and the projection of a textured 3D face model. To train such a n…

Cited by 80PDFScholar
2017

VQS: Linking Segmentations to Questions and Answers for Supervised Attention in VQA and Question-Focused Semantic Segmentation

ICCV 2017poster

Rich and dense human labeled datasets are the main enabling factor, among others, for the recent exciting work on vision-language understanding. Many seemingly distinct annotations (e.g., semantic segmentation and visual questions answering (VQA)) are inherently connected in that they reveal differe…

Cited by 145PDFcodeScholar
2016

A Multi-Level Contextual Model For Person Recognition in Photo Albums

CVPR 2016poster

In this work, we present a new framework for person recognition in photo albums that exploits contextual cues at multiple levels, spanning individual persons, individual photos, and photo groups. Through experiments, we show that the information available at each of these distinct contextual levels…

Cited by 39PDFScholar
2016

An egocentric computer vision based co-robot wheelchair

IROS 2016poster

Motivated by the emerging needs to improve the quality of life for the elderly and disabled individuals who rely on wheelchairs for mobility, and who might have limited or no hand functionality at all, we propose an egocentric computer vision based co-robot wheelchair to enhance their mobility witho…

Cited by 19SourceScholar
2015

A Convolutional Neural Network Cascade for Face Detection

CVPR 2015poster

In real-world face detection, large visual variations, such as those due to pose, expression, and lighting, demand an advanced discriminative model to accurately differentiate faces from the backgrounds. Consequently, effective models for the problem tend to be computationally prohibitive. To addre…

Cited by 1844SourcePDFScholar