← Search

Xuelin Chen

15 accepted papers

2026

EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered Decompositio

CVPR 2026

Existing video omnimatte methods typically rely on slow, multi-stage, or inference-time optimization pipelines that fail to fully exploit powerful generative priors, producing suboptimal decompositions. Our key insight is that, if a video inpainting model can be finetuned to remove the foreground-as

Cited by 0SourcecodeScholar
2026

LoST: Level of Semantics Tokenization for 3D Shapes

CVPR 2026

Tokenization is a fundamental technique in the generative modeling of various modalities. In particular, it plays a critical role in autoregressive (AR) models, which have recently emerged as a compelling option for 3D generation.However, optimal tokenization of 3D shapes remains an open question. S

Cited by 0SourcecodeScholar
2026

SpaceTimePilot: Generative Rendering of Dynamic Scenes Across Space and Time

CVPR 2026

We present SpaceTimePilot, a video diffusion model that disentangles space and time for controllable generative rendering. Given a monocular video, SpaceTimePilot can independently alter both the camera viewpoint and the motion sequence within the generative process, re-rendering the scene for conti

Cited by 0SourcecodeScholar
2026

V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties

CVPR 2026

Large-scale video generation models have shown remarkable potential in modeling photorealistic appearance and lighting interactions in real-world scenes. However, a closed-loop framework that jointly understands intrinsic scene properties (e.g., albedo, normal, material, and irradiance), leverages t

Cited by 0SourcecodeScholar
2025

CraftsMan3D: High-fidelity Mesh Generation with 3D Native Diffusion and Interactive Geometry Refiner

CVPR 2025poster

We present a novel generative 3D modeling system, coined CraftsMan, which can generate high-fidelity 3D geometries with highly varied shapes, regular mesh topologies, and detailed surfaces, and, notably, allows for refining the geometry in an interactive manner. Despite the significant advancements…

Cited by 0SourcePDFScholar
2025

DIDiffGes: Decoupled Semi-Implicit Diffusion Models for Real-time Gesture Generation from Speech

AAAI 2025technical

Diffusion models have demonstrated remarkable synthesis quality and diversity in generating co-speech gestures. However, the computationally intensive sampling steps associated with diffusion models hinder their practicality in real-world applications. Hence, we present DIDiffGes, for a Decoupled…

Cited by 0SourcePDFScholar
2025

Democratizing High-Fidelity Co-Speech Gesture Video Generation

ICCV 2025poster

Co-speech gesture video generation aims to synthesize realistic, audio-aligned videos of speakers, complete with synchronized facial expressions and body gestures. This task presents challenges due to the significant one-to-many mapping between audio and visual content, further complicated by the sc…

2024

SweetDreamer: Aligning Geometric Priors in 2D diffusion for Consistent Text-to-3D

ICLR 2024poster

It is inherently ambiguous to lift 2D results from pre-trained diffusion models to a 3D world for text-to-3D generation. 2D diffusion models solely learn view-agnostic priors and thus lack 3D knowledge during the lifting, leading to the multi-view inconsistency problem. We find that this problem pri…

2023

3D-Aware Object Goal Navigation via Simultaneous Exploration and Identification

CVPR 2023poster

Object goal navigation (ObjectNav) in unseen environments is a fundamental task for Embodied AI. Agents in existing works learn ObjectNav policies based on 2D maps, scene graphs, or image sequences. Considering this task happens in 3D space, a 3D-aware agent can advance its ObjectNav capability via…

Cited by 47SourcePDFScholar
2023

LivelySpeaker: Towards Semantic-Aware Co-Speech Gesture Generation

ICCV 2023poster

Gestures are non-verbal but important behaviors accompanying people's speech. While previous methods are able to generate speech rhythm-synchronized gestures, the semantic context of the speech is generally lacking in the gesticulations. Although semantic gestures do not occur very regularly in huma…

Cited by 26PDFcodeScholar
2023

Patch-Based 3D Natural Scene Generation From a Single Example

CVPR 2023poster

We target a 3D generative model for general natural scenes that are typically unique and intricate. Lacking the necessary volumes of training data, along with the difficulties of having ad hoc designs in presence of varying scene characteristics, renders existing setups intractable. Inspired by clas…

2022

VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated Objects

ICLR 2022poster

Perceiving and manipulating 3D articulated objects (e.g., cabinets, doors) in human environments is an important yet challenging task for future home-assistant robots. The space of 3D articulated objects is exceptionally rich in their myriad semantic categories, diverse shape geometry, and complicat…

Cited by 104SourcePDFScholar
2020

Multimodal Shape Completion via Conditional Generative Adversarial Networks

ECCV 2020poster

Several deep learning methods have been proposed for completing partial data from shape acquisition setups, i.e., filling the regions that were missing in the shape. These methods, however, only complete the partial shape with a single output, ignoring the ambiguity when reasoning the missing geomet…

2020

Unpaired Point Cloud Completion on Real Scans using Adversarial Training

ICLR 2020poster

As 3D scanning solutions become increasingly popular, several deep learning setups have been developed for the task of scan completion, i.e., plausibly filling in regions that were missed in the raw scans. These methods, however, largely rely on supervision in the form of paired training data, i.e.,…

Cited by 157SourcecodeScholar