← Search

Xiaohang Zhan

23 accepted papers

2026

UniSER: A Foundation Model for Unified Soft Effects Removal

CVPR 2026

Digital images are often degraded by soft effects such as lens flare, haze, shadows, and reflections, which reduce aesthetics even though the underlying pixels remain partially visible. The prevailing works address these degradations in isolation, developing highly specialized, specialist models tha

Cited by 0SourceScholar
2025

Stable-SCore: A Stable Registration-based Framework for 3D Shape Correspondence

CVPR 2025poster

Establishing character shape correspondence is a critical and fundamental task in computer vision and graphics, with diverse applications including re-topology, attribute transfer, and shape interpolation. Current dominant functional map methods, while effective in controlled scenarios, struggle in…

Cited by 0SourcePDFScholar
2025

Video-T1: Test-time Scaling for Video Generation

ICCV 2025poster

With the scale capability of increasing training data, model size, and computational cost, video generation has achieved impressive results in digital creation, enabling users to express creativity across various domains. Recently, researchers in Large Language Models (LLMs) have expanded the scalin…

Cited by 0SourcePDFScholar
2024

HumanGaussian: Text-Driven 3D Human Generation with Gaussian Splatting

CVPR 2024highlight

Realistic 3D human generation from text prompts is a desirable yet challenging task. Existing methods optimize 3D representations like mesh or neural fields via score distillation sampling (SDS) which suffers from inadequate fine details or excessive training time. In this paper we propose an effici…

Cited by 91SourcePDFScholar
2024

Programmable Motion Generation for Open-Set Motion Control Tasks

CVPR 2024highlight

Character animation in real-world scenarios necessitates a variety of constraints such as trajectories key-frames interactions etc. Existing methodologies typically treat single or a finite set of these constraint(s) as separate control tasks. These methods are often specialized and the tasks they a…

Cited by 5SourcePDFScholar
2024

TapMo: Shape-aware Motion Generation of Skeleton-free Characters

ICLR 2024poster

Previous motion generation methods are limited to the pre-rigged 3D human model, hindering their applications in the animation of various non-rigged characters. In this work, we present TapMo, a Text-driven Animation PIpeline for synthesizing Motion in a broad spectrum of skeleton-free 3D characters…

Cited by 11SourcePDFScholar
2023

Masked Frequency Modeling for Self-Supervised Visual Pre-Training

ICLR 2023poster

We present Masked Frequency Modeling (MFM), a unified frequency-domain-based approach for self-supervised pre-training of visual models. Instead of randomly inserting mask tokens to the input embeddings in the spatial domain, in this paper, we shift the perspective to the frequency domain. Specifica…

2023

RaBit: Parametric Modeling of 3D Biped Cartoon Characters With a Topological-Consistent Dataset

CVPR 2023poster

Assisting people in efficiently producing visually plausible 3D characters has always been a fundamental research topic in computer vision and computer graphics. Recent learning-based approaches have achieved unprecedented accuracy and efficiency in the area of 3D real human digitization. However, n…

Cited by 10SourcePDFScholar
2022

Style-ERD: Responsive and Coherent Online Motion Style Transfer

CVPR 2022poster

Motion style transfer is a common method for enriching character animation. Motion style transfer algorithms are often designed for offline settings where motions are processed in segments. However, for online animation applications, such as real-time avatar animation from motion capture, motions ne…

Cited by 33PDFScholar
2021

DetCo: Unsupervised Contrastive Learning for Object Detection

ICCV 2021poster

We present DetCo, a simple yet effective self-supervised approach for object detection. Unsupervised pre-training methods have been recently designed for object detection, but they are usually deficient in image classification, or the opposite. Unlike them, DetCo transfers well on downstream instanc…

Cited by 408PDFcodeScholar
2021

Unsupervised Object-Level Representation Learning from Scene Images

NeurIPS 2021poster

Contrastive self-supervised learning has largely narrowed the gap to supervised pre-training on ImageNet. However, its success highly relies on the object-centric priors of ImageNet, i.e., different augmented views of the same image correspond to the same object. Such a heavily curated constraint be…

2020

Exploiting Deep Generative Prior for Versatile Image Restoration and Manipulation

ECCV 2020poster

Learning a good image prior is a long-term goal for image restoration and manipulation. While existing methods like deep image prior (DIP) capture low-level image statistics, there are still gaps toward an image prior that captures rich image semantics including color, spatial coherence, textures, a…

2020

Learning to Cluster Faces via Confidence and Connectivity Estimation

CVPR 2020poster

Face clustering is an essential tool for exploiting the unlabeled face data, and has a wide range of applications including face annotation and retrieval. Recent works show that supervised clustering can result in noticeable performance gain. However, they usually involve heuristic steps and require…

Cited by 116PDFcodeScholar
2020

Online Deep Clustering for Unsupervised Representation Learning

CVPR 2020poster

Joint clustering and feature learning methods have shown remarkable performance in unsupervised representation learning. However, the training schedule alternating between feature clustering and network parameters update leads to unstable learning of visual representations. To overcome this challeng…

Cited by 255PDFcodeScholar
2019

Large-Scale Long-Tailed Recognition in an Open World

CVPR 2019oral

Real world data often have a long-tailed and open-ended distribution. A practical recognition system must classify among majority and minority classes, generalize from a few known instances, and acknowledge novelty upon a never seen instance. We define Open Long-Tailed Recognition (OLTR) as learning…

Cited by 1471PDFcodeScholar
2019

Learning to Cluster Faces on an Affinity Graph

CVPR 2019oral

Face recognition sees remarkable progress in recent years, and its performance has reached a very high level. Taking it to a next level requires substantially larger data, which would involve prohibitive annotation cost. Hence, exploiting unlabeled data becomes an appealing alternative. Recent works…

Cited by 161PDFcodeScholar
2019

Self-Supervised Learning via Conditional Motion Propagation

CVPR 2019poster

Intelligent agent naturally learns from motion. Various self-supervised algorithms have leveraged the motion cues to learn effective visual representations. The hurdle here is that motion is both ambiguous and complex, rendering previous works either suffer from degraded learning efficacy, or resort…

Cited by 63PDFcodeScholar
2019

Switchable Whitening for Deep Representation Learning

ICCV 2019poster

Normalization methods are essential components in convolutional neural networks (CNNs). They either standardize or whiten data using statistics estimated in predefined sets of pixels. Unlike existing works that design normalization techniques for specific tasks, we propose Switchable Whitening (SW),…

Cited by 193PDFcodeScholar
2018

Consensus-Driven Propagation in Massive Unlabeled Data for Face Recognition

ECCV 2018poster

Face recognition has witnessed great progresses in recent years, mainly attributed to the high-capacity model designed and the abundant labeled data collected. However, it becomes more and more prohibitive to scale up the current million-level identity annotations. In this work, we show that unlabel…