← Search

Xintong Han

27 accepted papers

2026

FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction

CVPR 2026

Current diffusion-based acceleration methods for long-portrait animation struggle to ensure identity (ID) consistency. This paper presents FlashPortrait, an end-to-end video diffusion transformer capable of synthesizing ID-preserving, infinite-length videos while achieving up to 6xacceleration in in

Cited by 0SourcecodeScholar
2026

MatPedia: A Universal Generative Foundation for High-Fidelity Material Synthesis

CVPR 2026

Physically-based rendering (PBR) materials are fundamental to photorealistic graphics, yet their creation remains labor-intensive and requires specialized expertise. While generative models have advanced material synthesis, existing methods lack a unified representation bridging natural image appear

Cited by 0SourceScholar
2025

MotionFollower: Editing Video Motion via Score-Guided Diffusion

ICCV 2025poster

Despite impressive advancements in diffusion-based video editing models in altering video attributes, there has been limited exploration into modifying motion information while preserving the original protagonist's appearance and background. In this paper, we propose MotionFollower, a score-guided d…

2025

StableAnimator: High-Quality Identity-Preserving Human Image Animation

CVPR 2025poster

Current diffusion models for human image animation struggle to ensure identity (ID) consistency. This paper presents StableAnimator, the first end-to-end ID-preserving video diffusion framework, which synthesizes high-quality videos without any post-processing, conditioned on a reference image and a…

2024

MotionEditor: Editing Video Motion via Content-Aware Diffusion

CVPR 2024poster

Existing diffusion-based video editing models have made gorgeous advances for editing attributes of a source video over time but struggle to manipulate the motion information while preserving the original protagonist's appearance and background. To address this we propose MotionEditor the first diff…

2024

PromptFusion: Decoupling Stability and Plasticity for Continual Learning

ECCV 2024poster

"Current research on continual learning mainly focuses on relieving catastrophic forgetting, and most of their success is at the cost of limiting the performance of newly incoming tasks. Such a trade-off is referred to as the stability-plasticity dilemma and is a more general and challenging problem…

2023

Human MotionFormer: Transferring Human Motions with Vision Transformers

ICLR 2023poster

Human motion transfer aims to transfer motions from a target dynamic person to a source static one for motion synthesis. An accurate matching between the source person and the target motion in both large and subtle motion changes is vital for improving the transferred motion quality. In this paper,…

2023

Multi-Prompt Alignment for Multi-Source Unsupervised Domain Adaptation

NeurIPS 2023poster

Most existing methods for unsupervised domain adaptation (UDA) rely on a shared network to extract domain-invariant features. However, when facing multiple source domains, optimizing such a network involves updating the parameters of the entire network, making it both computationally expensive and c…

2023

XFormer: Fast and Accurate Monocular 3D Body Capture

IJCAI 2023poster

We present XFormer, a novel human mesh and motion capture method that achieves real-time performance on consumer CPUs given only monocular images as input. The proposed network architecture contains two branches: a keypoint branch that estimates 3D human mesh vertices given 2D keypoints, and an imag…

Cited by 2SourcePDFScholar
2022

ObjectFormer for Image Manipulation Detection and Localization

CVPR 2022poster

Recent advances in image editing techniques have posed serious challenges to the trustworthiness of multimedia data, which drives the research of image tampering detection. In this paper, we propose ObjectFormer to detect and localize image manipulations. To capture subtle manipulation traces that a…

Cited by 190PDFScholar
2022

One Model to Edit Them All: Free-Form Text-Driven Image Manipulation with Semantic Modulations

NeurIPS 2022accept

Free-form text prompts allow users to describe their intentions during image manipulation conveniently. Based on the visual latent space of StyleGAN[21] and text embedding space of CLIP[34], studies focus on how to map these two latent spaces for text-driven attribute manipulations. Currently, the l…

2021

Action-guided 3D Human Motion Prediction

NeurIPS 2021poster

The ability of forecasting future human motion is important for human-machine interaction systems to understand human behaviors and make interaction. In this work, we focus on developing models to predict future human motion from past observed video frames. Motivated by the observation that human mo…

Cited by 10SourcePDFScholar
2021

DeFLOCNet: Deep Image Editing via Flexible Low-Level Controls

CVPR 2021poster

User-intended visual content fills the hole regions of an input image in the image editing scenario. The coarse lowlevel inputs, which typically consist of sparse sketch lines and color dots, convey user intentions for content creation (i.e., free-form editing). While existing methods combine an inp…

Cited by 42PDFcodeScholar
2021

Few-Shot Human Motion Transfer by Personalized Geometry and Texture Modeling

CVPR 2021poster

We present a new method for few-shot human motion transfer that achieves realistic human image generation with only a small number of appearance inputs. Despite recent advances in single person motion transfer, prior methods often require a large number of training images and take long training time…

Cited by 24PDFcodeScholar
2021

Fine-Grained Shape-Appearance Mutual Learning for Cloth-Changing Person Re-Identification

CVPR 2021poster

Recently, person re-identification (Re-ID) has achieved great progress. However, current methods largely depend on color appearance, which is not reliable when a person changes the clothes. Cloth-changing Re-ID is challenging since pedestrian images with clothes change exhibit large intra-class vari…

Cited by 206PDFScholar
2021

PD-GAN: Probabilistic Diverse GAN for Image Inpainting

CVPR 2021poster

We propose PD-GAN, a probabilistic diverse GAN forimage inpainting. Given an input image with arbitrary holeregions, PD-GAN produces multiple inpainting results withdiverse and visually realistic content. Our PD-GAN is builtupon a vanilla GAN which generates images based on random noise. During imag…

Cited by 287PDFcodeScholar
2019

FiNet: Compatible and Diverse Fashion Image Inpainting

ICCV 2019oral

Visual compatibility is critical for fashion analysis, yet is missing in existing fashion image synthesis systems. In this paper, we propose to explicitly model visual compatibility through fashion image inpainting. We present Fashion Inpainting Networks (FiNet), a two-stage image-to-image generatio…

Cited by 85PDFScholar
2019

Multi-Similarity Loss With General Pair Weighting for Deep Metric Learning

CVPR 2019poster

A family of loss functions built on pair-based computation have been proposed in the literature which provide a myriad of solutions for deep metric learning. In this pa-per, we provide a general weighting framework for under-standing recent pair-based loss functions. Our contributions are t…

Cited by 1000PDFcodeScholar
2018

DCAN: Dual Channel-wise Alignment Networks for Unsupervised Scene Adaptation

ECCV 2018poster

Harvesting dense pixel-level annotations to train deep neural networks for semantic segmentation is extremely expensive and unwieldy at scale. While learning from synthetic data where labels are readily available sounds promising, performance degrades significantly when testing on novel realistic da…

Cited by 317SourcePDFScholar
2018

Learning Rich Features for Image Manipulation Detection

CVPR 2018poster

Image manipulation detection is different from traditional semantic object detection because it pays more attention to tampering artifacts than to image content, which suggests that richer features need to be learned. We propose a two-stream Faster R-CNN network and train it end-to- end to detect th…

Cited by 797SourcePDFScholar
2018

NISP: Pruning Networks Using Neuron Importance Score Propagation

CVPR 2018poster

To reduce the significant redundancy in deep Convolutional Neural Networks (CNNs), most existing methods prune neurons by only considering the statistics of an individual layer or two consecutive layers (e.g., prune one layer to minimize the reconstruction error of the next layer), ignoring the effe…

Cited by 1103SourcePDFScholar
2017

Automatic Spatially-Aware Fashion Concept Discovery

ICCV 2017poster

This paper proposes an automatic spatially-aware concept discovery approach using weakly labeled image-text data from shopping websites. We first fine-tune GoogleNet by jointly modeling clothing images and their corresponding descriptions in a visual-semantic embedding space. Then, for each attribut…

Cited by 310PDFScholar
2017

Son of Zorn's lemma: Targeted style transfer using instance-aware semantic segmentation

ICASSP 2017accepted

Style transfer is an important task in which the style of a source image is mapped onto that of a target image. The method is useful for synthesizing derivative works of a particular artist or specific painting. This work considers targeted style transfer, in which the style of a template image is u…

Cited by 0SourceScholar
2015

Selecting Relevant Web Trained Concepts for Automated Event Retrieval

ICCV 2015poster

Complex event retrieval is a challenging research problem, especially when no training videos are available. An alternative to collecting training videos is to train a large semantic concept bank a priori. Given a text description of an event, event retrieval is performed by selecting concepts lingu…

Cited by 42PDFScholar