← Search

Zhuoyuan Li

14 accepted papers

2026

GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation

CVPR 2026

Open-vocabulary 3D semantic segmentation aims to segment arbitrary categories beyond the training set. Existing methods predominantly rely on distilling knowledge from 2D open-vocabulary models. However, aligning 3D features to the 2D representation space restricts intrinsic 3D geometric learning an

Cited by 0SourceScholar
2026

MeshSplat: Generalizable Sparse-View Surface Reconstruction via Gaussian Splatting

AAAI 2026technical

Surface reconstruction has been widely studied in computer vision and graphics. However, existing surface reconstruction works struggle to recover accurate scene geometry when the input views are extremely sparse. To address this issue, we propose MeshSplat, a generalizable sparse-view surface recon

Cited by 0SourcePDFScholar
2026

Perceptual Neural Video Compression with Color Separation and Rank Chain

CVPR 2026

Neural video compression (NVC) has achieved significant progress in recent years. The state-of-the-art (SOTA) NVC schemes, exemplified by the Deep Conditional Video Coding series, have focused on pursuing higher fidelity (e.g., PSNR), but lack sufficient exploitation of deep networks' advantages for

Cited by 0SourcecodeScholar
2026

ReFlow: Self-correction Motion Learning for Dynamic Scene Reconstruction

CVPR 2026

We present ReFlow, a unified framework for monocular dynamic scene reconstruction that learns 3D motion in a novel self-correction manner from raw video. Existing methods often suffer from incomplete scene initialization for dynamic regions, leading to unstable reconstruction and motion estimation,

Cited by 0SourceScholar
2025

BeyondMix: Leveraging Structural Priors and Long-Range Dependencies for Domain-Invariant LiDAR Segmentation

NeurIPS 2025poster

Domain adaptation for LiDAR semantic segmentation remains challenging due to the complex structural properties of point cloud data. While mix-based paradigms have shown promise, they often fail to fully leverage the rich structural priors inherent in 3D LiDAR point clouds. In this paper, we identify…

Cited by 0SourceScholar
2025

Pamba: Enhancing Global Interaction in Point Clouds via State Space Model

AAAI 2025technical

Transformers have demonstrated impressive results for 3D point cloud semantic segmentation. However, the quadratic complexity of transformer makes computation costs high, limiting the number of points that can be processed simultaneously and impeding the modeling of long-range dependencies between o…

Cited by 0SourcePDFScholar
2025

SAS: Segment Any 3D Scene with Integrated 2D Priors

ICCV 2025poster

The open vocabulary capability of 3D models is increasingly valued, as traditional methods with models trained with fixed categories fail to recognize unseen objects in complex dynamic 3D scenes. In this paper, we propose a simple yet effective approach, SAS, to integrate the open vocabulary capabil…

Cited by 0SourcePDFScholar
2025

Spend Wisely: Maximizing Post-Training Gains in Iterative Synthetic Data Bootstrapping

NeurIPS 2025spotlight

Modern foundation models often undergo iterative ``bootstrapping'' in their post-training phase: a model generates synthetic data, an external verifier filters out low-quality samples, and the high-quality subset is used for further fine-tuning. Over multiple iterations, the model performance improv…

Cited by 0SourceScholar
2024

M3DBench: Towards Omni 3D Assistant with Interleaved Multi-modal Instructions

ECCV 2024poster

"Recently, the understanding of the 3D world has garnered increased attention, facilitating autonomous agents to perform further decision-making. However, the majority of existing 3D vision-language datasets and methods are often limited to specific tasks, limiting their applicability in diverse sce…

Cited by 0SourcePDFScholar
2024

MotionChain: Conversational Motion Controllers via Multimodal Prompts

ECCV 2024poster

"Recent advancements in language models have demonstrated their adeptness in conducting multi-turn dialogues and retaining conversational context. However, this proficiency remains largely unexplored in other multimodal generative models, particularly in human motion models. By integrating multi-tur…

2024

Offline and Online Optical Flow Enhancement for Deep Video Compression

AAAI 2024technical

Video compression relies heavily on exploiting the temporal redundancy between video frames, which is usually achieved by estimating and using the motion information. The motion information is represented as optical flows in most of the existing deep video compression networks. Indeed, these network…

Cited by 20SourcePDFScholar