← Search

Xiaoxiao Long

34 accepted papers

2026

LottieGPT: Tokenizing Vector Animation for Autoregressive Generation

CVPR 2026

Despite rapid progress in video generation, existing models are incapable of producing vector animation, a dominant and highly expressive form of multimedia on the Internet. Vector animations offer resolution-independence, compactness, semantic structure, and editable parametric motion representatio

Cited by 0SourceScholar
2026

OccTENS: 3D Occupancy World Model Via Temporal Next-Scale Prediction

ICRA 2026poster

In this paper, we propose OccTENS, a generative occupancy world model that enables controllable, high-fidelity long-term occupancy generation while maintaining computational efficiency. Different from visual generation, the occupancy world model must capture the fine-grained 3D geometry and dynamic …

2026

OccTENS: 3D Occupancy World Model via Temporal Next-Scale Prediction

RA-L 2026

In this paper, we propose OccTENS, a generative occupancy world model that enables controllable, high-fidelity long-term occupancy generation while maintaining computational efficiency. Different from visual generation, the occupancy world model must capture the fine-grained 3D geometry and dynamic

Cited by 7SourceScholar
2026

Pressure2Motion: Hierarchical Human Motion Reconstruction from Ground Pressure with Text Guidance

CVPR 2026

We present Pressure2Motion, a novel motion capture algorithm that reconstructs human motion from a ground pressure sequence and text prompt. At inference time, Pressure2Motion requires only a pressure mat, eliminating the need for specialized lighting setups, cameras, or wearable devices, making it

Cited by 0SourcecodeScholar
2025

CADDreamer: CAD Object Generation from Single-view Images

CVPR 2025highlight

The field of diffusion-based 3D generation has experienced tremendous progress in recent times. However, existing 3D generative models often produce overly dense and unstructured meshes, which are in stark contrast to the compact, structured and clear-edged CAD models created by human modelers. We i…

Cited by 0SourcePDFScholar
2025

ComDrive: Comfort-Oriented End-to-End Autonomous Driving

IROS 2025

We propose ComDrive: the first comfort-oriented end-to-end autonomous driving system to generate temporally consistent and comfortable trajectories. Recent studies have demonstrated that imitation learning-based planners and learning-based trajectory scorers can effectively generate and select safet

Cited by 14SourcecodeScholar
2025

CraftsMan3D: High-fidelity Mesh Generation with 3D Native Diffusion and Interactive Geometry Refiner

CVPR 2025poster

We present a novel generative 3D modeling system, coined CraftsMan, which can generate high-fidelity 3D geometries with highly varied shapes, regular mesh topologies, and detailed surfaces, and, notably, allows for refining the geometry in an interactive manner. Despite the significant advancements…

Cited by 0SourcePDFScholar
2025

Dora: Sampling and Benchmarking for 3D Shape Variational Auto-Encoders

CVPR 2025poster

Recent 3D content generation pipelines commonly employ Variational Autoencoders (VAEs) to encode shapes into compact latent representations for diffusion-based generation. However, the widely adopted uniform point sampling strategy in Shape VAE training often leads to a significant loss of geometric…

2025

EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wild

CVPR 2025poster

Our work aims to reconstruct hand-object interactions from a single-view image, which is a fundamental but ill-posed task.Unlike methods that reconstruct from videos, multi-view images, or predefined 3D templates, single-view reconstruction faces significant challenges due to inherent ambiguities an…

2025

GGS: Generalizable Gaussian Splatting for Lane Switching in Autonomous Driving

AAAI 2025technical

We propose GGS, a Generalizable Gaussian Splatting method for Autonomous Driving that can achieve realistic rendering under large viewpoint changes. Previous generalizable 3D gaussian splatting methods are limited to rendering novel views that are very close to the original pair of images, which can…

Cited by 1SourcePDFScholar
2025

GoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving

CVPR 2025poster

We propose GoalFlow, an end-to-end autonomous driving method for generating high-quality multimodal trajectories. In autonomous driving scenarios, there is rarely a single suitable trajectory. Recent methods have increasingly focused on modeling multimodal trajectory distributions. However, they suf…

2025

MAGE : Single Image to Material-Aware 3D via the Multi-View G-Buffer Estimation Model

CVPR 2025poster

With advances in deep learning models and the availability of large-scale 3D datasets, we have recently witnessed significant progress in single-view 3D reconstruction. However, existing methods often fail to reconstruct physically based material properties given a single image, limiting their appli…

Cited by 0SourcePDFScholar
2025

OccRWKV: Rethinking Efficient 3D Semantic Occupancy Prediction with Linear Complexity

ICRA 2025

3D semantic occupancy prediction networks have demonstrated remarkable capabilities in reconstructing the geometric and semantic structure of 3D scenes, providing crucial information for robot navigation and autonomous driving systems. However, due to their large overhead from dense network structur

Cited by 10SourcecodeScholar
2024

DC-Gaussian: Improving 3D Gaussian Splatting for Reflective Dash Cam Videos

NeurIPS 2024poster

We present DC-Gaussian, a new method for generating novel views from in-vehicle dash cam videos. While neural rendering techniques have made significant strides in driving scenarios, existing methods are primarily designed for videos collected by autonomous vehicles. However, these videos are limite…

2024

Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise Attention

NeurIPS 2024poster

In this paper, we introduce **Era3D**, a novel multiview diffusion method that generates high-resolution multiview images from a single-view image. Despite significant advancements in multiview generation, existing methods still suffer from camera prior mismatch, inefficacy, and low resolution, resu…

Cited by 7SourcePDFScholar
2024

GaussianGrasper: 3D Language Gaussian Splatting for Open-Vocabulary Robotic Grasping

RA-L 2024

Constructing a 3D scene capable of accommodating open-ended language queries, is a pivotal pursuit in the domain of robotics, which facilitates robots in executing object manipulations based on human language directives. To achieve this, some research efforts have been dedicated to the development o

Cited by 102SourcecodeScholar
2024

GaussianPro: 3D Gaussian Splatting with Progressive Propagation

ICML 2024poster

3D Gaussian Splatting (3DGS) has recently revolutionized the field of neural rendering with its high fidelity and efficiency. However, 3DGS heavily depends on the initialized point cloud produced by Structure-from-Motion (SfM) techniques. When tackling large-scale scenes that unavoidably contain tex…

2024

GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image

ECCV 2024poster

"∗ Equal contributionWe introduce GeoWizard, a new generative foundation model designed for estimating geometric attributes, , depth and normals, from single images. While significant research has already been conducted in this area, the progress has been substantially limited by the low diversity a…

Cited by 103SourcePDFScholar
2024

LiveHPS++: Robust and Coherent Motion Capture in Dynamic Free Environment

ECCV 2024oral

"LiDAR-based human motion capture has garnered significant interest in recent years for its practicability in large-scale and unconstrained environments. However, most methods rely on cleanly segmented human point clouds as input, the accuracy and smoothness of their motion results are compromised w…

2024

MonoOcc: Digging into Monocular Semantic Occupancy Prediction

ICRA 2024poster

Monocular Semantic Occupancy Prediction aims to infer the complete 3D geometry and semantic information of scenes from only 2D images. It has garnered significant attention, particularly due to its potential to enhance the 3D perception of autonomous vehicles. However, existing methods rely on a com…

Cited by 31SourcecodeScholar
2024

Surf-D: Generating High-Quality Surfaces of Arbitrary Topologies Using Diffusion Models

ECCV 2024poster

"We present Surf-D, a novel method for generating high-quality 3D shapes as Surfaces with arbitrary topologies using Diffusion models. Previous methods explored shape generation with different representations and they suffer from limited topologies and poor geometry details. To generate high-quality…

Cited by 1SourcePDFScholar
2024

SyncDreamer: Generating Multiview-consistent Images from a Single-view Image

ICLR 2024spotlight

In this paper, we present a novel diffusion model called SyncDreamer that generates multiview-consistent images from a single-view image. Using pretrained large-scale 2D diffusion models, recent work Zero123 demonstrates the ability to generate plausible novel views from a single-view image of an ob…

2024

TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes

ECCV 2024poster

"3D dense captioning stands as a cornerstone in achieving a comprehensive understanding of 3D scenes through natural language. It has recently witnessed remarkable achievements, particularly in indoor settings. However, the exploration of 3D dense captioning in outdoor scenes is hindered by two majo…

2024

UC-NERF: Neural Radiance Field for Under-Calibrated Multi-View Cameras in Autonomous Driving

ICLR 2024poster

Multi-camera setups find widespread use across various applications, such as autonomous driving, as they greatly expand sensing capabilities. Despite the fast development of Neural radiance field (NeRF) techniques and their wide applications in both indoor and outdoor scenes, applying NeRF to multi…

Cited by 9SourcePDFScholar
2024

Wonder3D: Single Image to 3D using Cross-Domain Diffusion

CVPR 2024highlight

In this work we introduce Wonder3D a novel method for generating high-fidelity textured meshes from single-view images with remarkable efficiency. Recent methods based on the Score Distillation Sampling (SDS) loss methods have shown the potential to recover 3D geometry from 2D diffusion priors but t…

Cited by 414SourcePDFScholar
2023

Learning Long-Range Information with Dual-Scale Transformers for Indoor Scene Completion

ICCV 2023poster

Due to the limited resolution of 3D sensors and the inevitable mutual occlusion between objects, 3D scans of real scenes are commonly incomplete. Previous scene completion methods struggle to capture long-range spatial feature, resulting in unsatisfactory completion results. To alleviate the pro…

Cited by 3PDFScholar
2023

NeTO:Neural Reconstruction of Transparent Objects with Self-Occlusion Aware Refraction-Tracing

ICCV 2023poster

We present a novel method called NeTO, for capturing the 3D geometry of solid transparent objects from 2D images via volume rendering. Reconstructing transparent objects is a very challenging task, which is ill-suited for general-purpose reconstruction techniques due to the specular light transport…

Cited by 12PDFScholar
2023

NeuralUDF: Learning Unsigned Distance Fields for Multi-View Reconstruction of Surfaces With Arbitrary Topologies

CVPR 2023poster

We present a novel method, called NeuralUDF, for reconstructing surfaces with arbitrary topologies from 2D images via volume rendering. Recent advances in neural rendering based reconstruction have achieved compelling results. However, these methods are limited to objects with closed surfaces since…

Cited by 67SourcePDFScholar
2022

Gen6D: Generalizable Model-Free 6-DoF Object Pose Estimation from RGB Images

ECCV 2022poster

"In this paper, we present a generalizable model-free 6-DoF object pose estimator called Gen6D. Existing generalizable pose estimators either need the high-quality object models or require additional depth maps or object masks in test time, which significantly limits their application scope. In cont…

2022

NeuRIS: Neural Reconstruction of Indoor Scenes Using Normal Priors

ECCV 2022poster

"Reconstructing 3D indoor scenes from 2D images is an important task in many computer vision and graphics applications. A main challenge in this task is that large texture-less areas in typical indoor scenes make existing methods struggle to produce satisfactory reconstruction results. We propose a…

Cited by 113SourcePDFScholar
2022

SparseNeuS: Fast Generalizable Neural Surface Reconstruction from Sparse Views

ECCV 2022poster

"We introduce SparseNeuS, a novel neural rendering based method for the task of surface reconstruction from multi-view images. This task becomes more difficult when only sparse images are provided as input, a scenario where existing neural reconstruction approaches usually produce incomplete or dist…

Cited by 195SourcePDFScholar
2021

Adaptive Surface Normal Constraint for Depth Estimation

ICCV 2021poster

We present a novel method for single image depth estimation using surface normal constraints. Existing depth estimation methods either suffer from the lack of geometric constraints, or are limited to the difficulty of reliably capturing geometric context, which leads to a bottleneck of depth estimat…

Cited by 72PDFcodeScholar
2021

Multi-view Depth Estimation using Epipolar Spatio-Temporal Networks

CVPR 2021poster

We present a novel method for multi-view depth estimation from a single video, which is a critical task in various applications, such as perception, reconstruction and robot navigation. Although previous learning-based methods have demonstrated compelling results, most works estimate depth maps of i…

Cited by 88PDFcodeScholar
2020

Occlusion-Aware Depth Estimation with Adaptive Normal Constraints

ECCV 2020poster

We present a new learning-based method for multi-frame depth estimation from a color video, which is a fundamental problem in scene understanding, robot navigation or handheld 3D reconstruction. While recent learning-based methods estimate depth at high accuracy, 3D point clouds exported from their…