← Search

Xiaohui Xie

20 accepted papers

2026

CoMA: Compositional Human Motion Generation with Multi-modal Agents

AAAI 2026technical

3D human motion generation has seen substantial advancement in recent years. While state-of-the-art approaches have improved performance significantly, they still struggle with complex and detailed motions unseen in training data, largely due to the scarcity of motion datasets and the prohibitive co

Cited by 0SourcePDFScholar
2026

CoRA: A Collaborative Robust Architecture with Hybrid Fusion for Efficient Perception

AAAI 2026technical

Collaborative perception has garnered significant attention as a crucial technology to overcome the perceptual limitations of single-agent systems. Many state-of-the-art (SOTA) methods have achieved communication efficiency and high performance via intermediate fusion. However, they share a critical

Cited by 0SourcePDFScholar
2026

OPLoRA: Orthogonal Projection LoRA Prevents Catastrophic Forgetting During Parameter-Efficient Fine-Tuning

AAAI 2026technical

Low-Rank Adaptation (LoRA) enables efficient fine-tuning of large language models but suffers from catastrophic forgetting when learned updates interfere with the dominant singular directions that encode essential pre-trained knowledge. We propose Orthogonal Projection LoRA (OPLoRA), a theoretically

Cited by 0SourcePDFScholar
2025

Ouroboros: Single-step Diffusion Models for Cycle-consistent Forward and Inverse Rendering

ICCV 2025poster

While multi-step diffusion models have advanced both forward and inverse rendering, existing approaches often treat these problems independently, leading to cycle inconsistency and slow inference speed. In this work, we present Ouroboros, a framework composed of two single-step diffusion models that…

Cited by 0SourcePDFScholar
2025

UMAMI: Unifying Masked Autoregressive Models and Deterministic Rendering for View Synthesis

NeurIPS 2025poster

Novel view synthesis (NVS) seeks to render photorealistic, 3D‑consistent images of a scene from unseen camera poses given only a sparse set of posed views. Existing deterministic networks render observed regions quickly but blur unobserved areas, whereas stochastic diffusion‑based methods hallucinat…

Cited by 0SourceScholar
2024

Diffeomorphic Mesh Deformation via Efficient Optimal Transport for Cortical Surface Reconstruction

ICLR 2024poster

Mesh deformation plays a pivotal role in many 3D vision tasks including dynamic simulations, rendering, and reconstruction. However, defining an efficient discrepancy between predicted and target meshes remains an open problem. A prevalent approach in current deep learning is the set-based approach…

Cited by 1SourcePDFScholar
2024

Integrating Efficient Optimal Transport and Functional Maps For Unsupervised Shape Correspondence Learning

CVPR 2024poster

In the realm of computer vision and graphics accurately establishing correspondences between geometric 3D shapes is pivotal for applications like object tracking registration texture transfer and statistical shape analysis. Moving beyond traditional hand-crafted and data-driven feature learning meth…

Cited by 4SourcePDFScholar
2024

LidaRF: Delving into Lidar for Neural Radiance Field on Street Scenes

CVPR 2024highlight

Photorealistic simulation plays a crucial role in applications such as autonomous driving where advances in neural radiance fields (NeRFs) may allow better scalability through the automatic creation of digital 3D assets. However reconstruction quality suffers on street scenes due to largely collinea…

Cited by 2SourcePDFScholar
2024

MaskINT: Video Editing via Interpolative Non-autoregressive Masked Transformers

CVPR 2024poster

Recent advances in generative AI have significantly enhanced image and video editing particularly in the context of text prompt control. State-of-the-art approaches predominantly rely on diffusion models to accomplish these tasks. However the computational demands of diffusion-based methods are subs…

Cited by 4SourcePDFScholar
2023

Peeling the Onion: Hierarchical Reduction of Data Redundancy for Efficient Vision Transformer Training

AAAI 2023technical

Vision transformers (ViTs) have recently obtained success in many applications, but their intensive computation and heavy memory usage at both training and inference time limit their generalization. Previous compression algorithms usually start from the pre-trained dense models and only focus on eff…

2022

EI-CLIP: Entity-Aware Interventional Contrastive Learning for E-Commerce Cross-Modal Retrieval

CVPR 2022poster

recommendation, and marketing services. Extensive efforts have been made to conquer the cross-modal retrieval problem in the general domain. When it comes to E-commerce, a common practice is to adopt the pretrained model and finetune on E-commerce data. Despite its simplicity, the performance is sub…

Cited by 68PDFScholar
2022

Identity-Aware Hand Mesh Estimation and Personalization from RGB Images

ECCV 2022poster

"Reconstructing 3D hand meshes from monocular RGB images has attracted increasing amount of attention due to its enormous potential applications in the field of AR/VR. Most state-of-the-art methods attempt to tackle this task in an anonymous manner. Specifically, the identity of the subject is ignor…

2022

PPT: Token-Pruned Pose Transformer for Monocular and Multi-View Human Pose Estimation

ECCV 2022poster

"Recently, the vision transformer and its variants have played an increasingly important role in both monocular and multi-view human pose estimation. Considering image patches as tokens, transformers can model the global dependencies within the entire image or across images from other views. However…

2022

Topology-Preserving Shape Reconstruction and Registration via Neural Diffeomorphic Flow

CVPR 2022poster

Deep Implicit Functions (DIFs) represent 3D geometry with continuous signed distance functions learned through deep neural nets. Recently DIFs-based methods have been proposed to handle shape reconstruction and dense point correspondences simultaneously, capturing semantic relationships across shape…

Cited by 45PDFcodeScholar
2021

Recurrent Mask Refinement for Few-Shot Medical Image Segmentation

ICCV 2021poster

Although having achieved great success in medical image segmentation, deep convolutional neural networks usually require a large dataset with manual annotations for training and are difficult to generalize to unseen classes. Few-shot learning has the potential to address these challenges by learning…

Cited by 151PDFcodeScholar
2021

Test-Time Training for Deformable Multi-Scale Image Registration

ICRA 2021poster

Registration is a fundamental task in medical robotics and is often a crucial step for many downstream tasks such as motion analysis, intra-operative tracking and image segmentation. Popular registration methods such as ANTs and NiftyReg optimize objective functions for each pair of images from scra…

Cited by 32SourceScholar
2021

Undistillable: Making A Nasty Teacher That CANNOT teach students

ICLR 2021spotlight

Knowledge Distillation (KD) is a widely used technique to transfer knowledge from pre-trained teacher models to (usually more lightweight) student models. However, in certain situations, this technique is more of a curse than a blessing. For instance, KD poses a potential risk of exposing intellect…

2020

Dynamically Pruned Message Passing Networks for Large-scale Knowledge Graph Reasoning

ICLR 2020poster

We propose Dynamically Pruned Message Passing Networks (DPMPN) for large-scale knowledge graph reasoning. In contrast to existing models, embedding-based or path-based, we learn an input-dependent subgraph to explicitly model a sequential reasoning process. Each subgraph is dynamically constructed,…

Cited by 86SourcecodeScholar
2019

VTNFP: An Image-Based Virtual Try-On Network With Body and Clothing Feature Preservation

ICCV 2019poster

Image-based virtual try-on systems with the goal of transferring a desired clothing item onto the corresponding region of a person have made great strides recently, but challenges remain in generating realistic looking images that preserve both body and clothing details. Here we present a new virtua…

Cited by 203PDFScholar