← Search

Yiming Xie

13 accepted papers

2026

LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction

CVPR 2026

Recent feed-forward reconstruction models like VGGT and \pi^3 achieve impressive reconstruction quality but cannot process streaming videos due to quadratic memory complexity, limiting their practical deployment. While existing streaming methods address this through learned memory mechanisms or caus

Cited by 0SourcecodeScholar
2025

Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression

CVPR 2025poster

Since 2023, Vector Quantization (VQ)-based discrete generation methods have rapidly dominated human motion generation, primarily surpassing diffusion-based continuous generation methods in standard performance metrics. However, VQ-based methods have inherent limitations. Representing continuous moti…

Cited by 0SourcePDFScholar
2025

SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation

ICCV 2025poster

We present Stable Video 4D 2.0 (SV4D 2.0), a multi-view video diffusion model for dynamic 3D asset generation. Compared to its predecessor SV4D, SV4D 2.0 is more robust to occlusions and large motion, generalizes better to real-world videos, and produces higher-quality outputs in terms of detail sha…

Cited by 0SourcePDFScholar
2025

SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency

ICLR 2025poster

We present Stable Video 4D (SV4D) — a latent video diffusion model for multi-frame and multi-view consistent dynamic 3D content generation. Unlike previous methods that rely on separately trained generative models for video generation and novel view synthesis, we design a unified diffusion model to…

2025

Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs

NeurIPS 2025poster

Unlocking spatial reasoning in Multimodal Large Language Models (MLLMs) is crucial for enabling intelligent interaction with 3D environments. While prior efforts often rely on explicit 3D inputs or specialized model architectures, we ask: can MLLMs reason about 3D space using only structured 2D repr…

Cited by 0SourceScholar
2024

OmniControl: Control Any Joint at Any Time for Human Motion Generation

ICLR 2024poster

We present a novel approach named OmniControl for incorporating flexible spatial control signals into a text-conditioned human motion generation model based on the diffusion process. Unlike previous methods that can only control the pelvis trajectory, OmniControl can incorporate flexible spatial con…

2024

SynFog: A Photo-realistic Synthetic Fog Dataset based on End-to-end Imaging Simulation for Advancing Real-World Defogging in Autonomous Driving

CVPR 2024poster

To advance research in learning-based defogging algorithms various synthetic fog datasets have been developed. However exsiting datasets created using the Atmospheric Scattering Model (ASM) or real-time rendering engines often struggle to produce photo-realistic foggy images that accurately mimic th…

Cited by 4SourcePDFScholar
2023

Pixel-Aligned Recurrent Queries for Multi-View 3D Object Detection

ICCV 2023poster

We present PARQ - a multi-view 3D object detector with transformer and pixel-aligned recurrent queries. Unlike previous works that use learnable features or only encode 3D point positions as queries in the decoder, PARQ leverages appearance-enhanced queries initialized from reference points in 3D sp…

Cited by 9PDFcodeScholar
2022

PlanarRecon: Real-Time 3D Plane Detection and Reconstruction From Posed Monocular Videos

CVPR 2022poster

We present PlanarRecon -- a novel framework for globally coherent detection and reconstruction of 3D planes from a posed monocular video. Unlike previous works that detect planes in 2D from a single image, PlanarRecon incrementally detects planes in 3D for each video fragment, which consists of a se…

Cited by 30PDFcodeScholar
2021

NeuralRecon: Real-Time Coherent 3D Reconstruction From Monocular Video

CVPR 2021poster

We present a novel framework named NeuralRecon for real-time 3D scene reconstruction from a monocular video. Unlike previous methods that estimate single-view depth maps separately on each key-frame and fuse them later, we propose to directly reconstruct local surfaces represented as sparse TSDF vol…

Cited by 358PDFcodeScholar
2021

You Don't Only Look Once: Constructing Spatial-Temporal Memory for Integrated 3D Object Detection and Tracking

ICCV 2021poster

Humans are able to continuously detect and track surrounding objects by constructing a spatial-temporal memory of the objects when looking around. In contrast, 3D object detectors in existing tracking-by-detection systems often search for objects in every new video frame from scratch, without fully…

Cited by 13PDFcodeScholar
2020

Disp R-CNN: Stereo 3D Object Detection via Shape Prior Guided Instance Disparity Estimation

CVPR 2020poster

In this paper, we propose a novel system named Disp R-CNN for 3D object detection from stereo images. Many recent works solve this problem by first recovering a point cloud with disparity estimation and then apply a 3D detector. The disparity map is computed for the entire image, which is costly and…

Cited by 145PDFcodeScholar