← Search

Haozhe Xie

13 accepted papers

2025

3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion

CVPR 2025highlight

The increasing demand for high-quality 3D assets across various industries necessitates efficient and automated 3D content creation. Despite recent advancements in 3D generative models, existing methods still face challenges with optimization speed, geometric fidelity, and the lack of assets for phy…

2025

DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes

ICLR 2025spotlight

Urban scene generation has been developing rapidly recently. However, existing methods primarily focus on generating static and single-frame scenes, overlooking the inherently dynamic nature of real-world driving environments. In this work, we introduce DynamicCity, a novel 4D occupancy generation f…

Cited by 0SourcePDFScholar
2025

Multi-view Consistent 3D Panoptic Scene Understanding

AAAI 2025technical

3D panoptic scene understanding seeks to create novel view images with 3D-consistent panoptic segmentation, which is crucial for many vision and robotics applications. Mainstream methods (e.g., Panoptic Lifting) directly use machine-generated 2D panoptic segmentation masks as training labels. Howeve…

Cited by 0SourcePDFScholar
2024

Blur-aware Spatio-temporal Sparse Transformer for Video Deblurring

CVPR 2024poster

Video deblurring relies on leveraging information from other frames in the video sequence to restore the blurred regions in the current frame. Mainstream approaches employ bidirectional feature propagation spatio-temporal transformers or a combination of both to extract information from the video se…

2024

CityDreamer: Compositional Generative Model of Unbounded 3D Cities

CVPR 2024poster

3D city generation is a desirable yet challenging task since humans are more sensitive to structural distortions in urban environments. Additionally generating 3D cities is more complex than 3D natural scenes since buildings as objects of the same class exhibit a wider range of appearances compared…

2022

Spatio-Temporal Deformable Attention Network for Video Deblurring

ECCV 2022poster

"The key success factor of the video deblurring methods is to compensate for the blurry pixels of the mid-frame with the sharp pixels of the adjacent video frames. Therefore, mainstream methods align the adjacent frames based on the estimated optical flows and fuse the alignment frames for restorati…

2021

Efficient Regional Memory Network for Video Object Segmentation

CVPR 2021poster

Recently, several Space-Time Memory based networks have shown that the object cues (e.g. video frames as well as the segmented object masks) from the past frames are useful for segmenting objects in the current frame. However, these methods exploit the information from the memory by global-to-global…

Cited by 188PDFcodeScholar
2020

GRNet: Gridding Residual Network for Dense Point Cloud Completion

ECCV 2020poster

Estimating the complete 3D point cloud from an incomplete one is a key problem in many vision and robotics applications. Mainstream methods (e.g., PCN and TopNet) use Multi-layer Perceptrons (MLPs) to directly process point clouds, which may cause the loss of details because the structural and conte…

2019

DAVANet: Stereo Deblurring With View Aggregation

CVPR 2019oral

Nowadays stereo cameras are more commonly adopted in emerging devices such as dual-lens smartphones and unmanned aerial vehicles. However, they also suffer from blurry images in dynamic scenes which leads to visual discomfort and hampers further image processing. Previous works have succeeded in mon…

Cited by 119PDFScholar
2019

Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View Images

ICCV 2019poster

Recovering the 3D representation of an object from single-view or multi-view RGB images by deep neural networks has attracted increasing attention in the past few years. Several mainstream works (e.g., 3D-R2N2) use recurrent neural networks (RNNs) to fuse multiple feature maps extracted from input i…

Cited by 471PDFcodeScholar
2019

Spatio-Temporal Filter Adaptive Network for Video Deblurring

ICCV 2019poster

Video deblurring is a challenging task due to the spatially variant blur caused by camera shake, object motions, and depth variations, etc. Existing methods usually estimate optical flow in the blurry video to align consecutive frames or approximate blur kernels. However, they tend to generate artif…

Cited by 247PDFScholar