← Search

Hu Han

13 accepted papers

2025

EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion Models

ICCV 2025poster

The progress on generative models has led to significant advances on text-to-video (T2V) generation, yet the motion controllability of generated videos remains limited. Existing motion transfer approaches explored the motion representations of reference videos to guide generation. Nevertheless, thes…

2024

Decoupled Textual Embeddings for Customized Image Generation

AAAI 2024technical

Customized text-to-image generation, which aims to learn user-specified concepts with a few images, has drawn significant attention recently. However, existing methods usually suffer from overfitting issues and entangle the subject-unrelated information (e.g., background and pose) with the learned c…

2023

DISC: Learning From Noisy Labels via Dynamic Instance-Specific Selection and Correction

CVPR 2023poster

Existing studies indicate that deep neural networks (DNNs) can eventually memorize the label noise. We observe that the memorization strength of DNNs towards each instance is different and can be represented by the confidence value, which becomes larger and larger during the training process. Based…

2023

Data-Free Knowledge Distillation via Feature Exchange and Activation Region Constraint

CVPR 2023poster

Despite the tremendous progress on data-free knowledge distillation (DFKD) based on synthetic data generation, there are still limitations in diverse and efficient data synthesis. It is naive to expect that a simple combination of generative network-based data synthesis and data augmentation will so…

2023

Modeling the Relative Visual Tempo for Self-supervised Skeleton-based Action Recognition

ICCV 2023poster

Visual tempo characterizes the dynamics and the temporal evolution, which helps describe actions. Recent approaches directly perform visual tempo prediction on skeleton sequences, which may suffer from insufficient feature representation issue. In this paper, we observe that relative visual tempo is…

Cited by 25PDFcodeScholar
2022

Towards High-Fidelity Face Self-Occlusion Recovery via Multi-View Residual-Based GAN Inversion

AAAI 2022technical

Face self-occlusions are inevitable due to the 3D nature of the human face and the loss of information in the projection process from 3D to 2D images. While recovering face self-occlusions based on 3D face reconstruction, e.g., 3D Morphable Model (3DMM) and its variants provides an effective solutio…

Cited by 7SourcePDFScholar
2020

Cross-Domain Face Presentation Attack Detection via Multi-Domain Disentangled Representation Learning

CVPR 2020poster

Face presentation attack detection (PAD) has been an urgent problem to be solved in the face recognition systems. Conventional approaches usually assume the testing and training are within the same domain; as a result, they may not generalize well into unseen scenarios because the representations le…

Cited by 234PDFcodeScholar
2020

Video-based Remote Physiological Measurement via Cross-verified Feature Disentangling

ECCV 2020poster

Remote physiological measurements, e.g., remote photoplethysmography (rPPG) based heart rate (HR), heart rate variability (HRV) and respiration frequency (RF) measuring, are playing more and more important roles under the application scenarios where contact measurement is inconvenient or impossible.…

2019

Local Relationship Learning With Person-Specific Shape Regularization for Facial Action Unit Detection

CVPR 2019poster

Encoding individual facial expressions via action units (AUs) coded by the Facial Action Coding System (FACS) has been found to be an effective approach in resolving the ambiguity issue among different expressions. While a number of methods have been proposed for AU detection, robust AU detection in…

Cited by 171PDFScholar
2019

Multi-label Co-regularization for Semi-supervised Facial Action Unit Recognition

NeurIPS 2019poster

Facial action units (AUs) recognition is essential for emotion analysis and has been widely applied in mental state analysis. Existing work on AU recognition usually requires big face dataset with accurate AU labels. However, manual AU annotation requires expertise and can be time-consuming. In this…