← Search

Xudong XU

16 accepted papers

2026

ILR-SMO: Iterative Latent Refinement for Robust Spatial Multi-Omics Integration

IJCAI 2026

Spatial multi-omics technologies jointly profile diverse molecular modalities with spatial context, providing a comprehensive view of cellular heterogeneity and tissue organization. To integrate spatial multi-omics data and identify spatial domains, a wide range of unsupervised methods has been prop

Cited by 0Scholar
2026

STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics–Physics Dual System

ICML 2026poster

Generating simulation-ready tabletop scenes from task instructions is an intriguing and promising research direction in the field of Embodied AI. However, existing task-to-scene generation methods rely exclusively on large language models (LLMs) to predict scene layouts, inevitably yielding object c…

Cited by 0SourceScholar
2025

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts

NeurIPS 2025poster

The advancement of Embodied AI heavily relies on large-scale, simulatable 3D scene datasets characterized by scene diversity and realistic layouts. However, existing datasets typically suffer from limitations in data scale or diversity, sanitized layouts lacking small items, and severe object collis…

Cited by 0SourceScholar
2025

MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning

NeurIPS 2025spotlight

The ability of robots to interpret human instructions and execute manipulation tasks necessitates the availability of task-relevant tabletop scenes for training. However, traditional methods for creating these scenes rely on time-consuming manual layout design or purely randomized layouts, which are…

Cited by 0SourceScholar
2025

MeshCoder: LLM-Powered Structured Mesh Code Generation from Point Clouds

NeurIPS 2025poster

Reconstructing 3D objects into editable programs is pivotal for applications like reverse engineering and shape editing. However, existing methods often rely on limited domain-specific languages (DSLs) and small-scale datasets, restricting their ability to model complex geometries and structures. To…

Cited by 0SourceScholar
2024

RoomTex: Texturing Compositional Indoor Scenes via Iterative Inpainting

ECCV 2024poster

"The advancement of diffusion models has pushed the boundary of text-to-3D object generation. While it is straightforward to composite objects into a scene with reasonable geometry, it is nontrivial to texture such a scene perfectly due to style inconsistency and occlusions between objects. To tackl…

2024

Text to Layer-wise 3D Clothed Human Generation

ECCV 2024poster

"This paper addresses the task of 3D clothed human generation from textural descriptions. Previous works usually encode the human body and clothes as a holistic model and generate the whole model in a single-stage optimization, which makes them struggle for clothing editing and meanwhile lose fine-g…

Cited by 12SourcePDFScholar
2023

Sounding Bodies: Modeling 3D Spatial Sound of Humans Using Body Pose and Audio

NeurIPS 2023spotlight

While 3D human body modeling has received much attention in computer vision, modeling the acoustic equivalent, i.e. modeling 3D spatial audio produced by body motion and speech, has fallen short in the community. To close this gap, we present a model that can generate accurate 3D spatial audio for f…

2023

Voxurf: Voxel-based Efficient and Accurate Neural Surface Reconstruction

ICLR 2023top-25%

Neural surface reconstruction aims to reconstruct accurate 3D surfaces based on multi-view images. Previous methods based on neural volume rendering mostly train a fully implicit model with MLPs, which typically require hours of training for a single scene. Recent efforts explore the explicit volume…

2022

A Conditional Point Diffusion-Refinement Paradigm for 3D Point Cloud Completion

ICLR 2022poster

3D point clouds are an important data format that captures 3D information for real world objects. Since 3D point clouds scanned in the real world are often incomplete, it is important to recover the complete point cloud for many downstreaming applications. Most existing point cloud completion metho…

2021

A Shading-Guided Generative Implicit Model for Shape-Accurate 3D-Aware Image Synthesis

NeurIPS 2021poster

The advancement of generative radiance fields has pushed the boundary of 3D-aware image synthesis. Motivated by the observation that a 3D object should look realistic from multiple viewpoints, these methods introduce a multi-view constraint as regularization to learn valid 3D radiance fields from 2D…

2021

Generative Occupancy Fields for 3D Surface-Aware Image Synthesis

NeurIPS 2021poster

The advent of generative radiance fields has significantly promoted the development of 3D-aware image synthesis. The cumulative rendering process in radiance fields makes training these generative models much easier since gradients are distributed over the entire volume, but leads to diffused object…

2021

Visually Informed Binaural Audio Generation without Binaural Audios

CVPR 2021poster

Stereophonic audio, especially binaural audio, plays an essential role in immersive viewing environments. Recent research has explored generating stereophonic audios guided by visual cues and multi-channel audio collections in a fully-supervised manner. However, due to the requirement of professiona…

Cited by 65PDFScholar
2020

Sep-Stereo: Visually Guided Stereophonic Audio Generation by Associating Source Separation

ECCV 2020poster

Stereophonic audio is an indispensable ingredient to enhance human auditory experience. Recent research has explored the usage of visual information as guidance to generate binaural or ambisonic audio from mono ones with stereo supervision. However, this fully supervised paradigm suffers from an inh…

Cited by 103SourcePDFScholar