CVPC: Cross-Modal Visual-Guided Point Cloud Completion
Bhanu Pratap Paregi, Vaibhav Kumar
Abstract
Robust 3D scene understanding and mapping is criticalto various applications such as autonomous systems and 3D scene development. The perception engine in these applications face challenges related to incomplete point clouds captured through sensors in unstructured environments. Existing models rely upon per sample image and point cloud pair in the completion task, limiting its usefulness in fast processing and accurate outcomes. To overcome the limitations we present Cross-Modal Visual-Guided Point Cloud Completion (CVPC), a model that fuses geometric transformers with category-level multiview visual priors to avoid per-sample image alignment. CVPC stores one diffusion-derived prototype per category and injects semantic cues through cross-modal attention to recover edges and thin structures. The alignment-free design supports real-time deployment on embedded platforms. CVPC achieves state-of-the-art results on shapenet-55 (<inline-formula><tex-math notation="LaTeX">$0.73\times 10^{-3}$</tex-math></inline-formula> Chamfer Distance-L2, <inline-formula><tex-math notation="LaTeX">$\text{CD}\text{-}\text{L2}$</tex-math></inline-formula>), PCN manipulation objects (<inline-formula><tex-math notation="LaTeX">$6.32\times 10^{-3}$</tex-math></inline-formula> Chamfer Distance-L1, <inline-formula><tex-math notation="LaTeX">$\text{CD}\text{-}\text{L1}$</tex-math></inline-formula>), and KITTI-Cars (0.388 Minimum Matching Distance, <inline-formula><tex-math notation="LaTeX">$\text{MMD}$</tex-math></inline-formula>). Qualitative results demonstrate improved completion of chair legs, car roofs, and lamp stems relative to geometry-only baselines.
BibTeX
@inproceedings{ral2026_cvpccrossmodalvi,
title = {CVPC: Cross-Modal Visual-Guided Point Cloud Completion},
author = {Bhanu Pratap Paregi and Vaibhav Kumar},
booktitle = {RA-L 2026},
year = {2026}
}