← Search

Cheng Sun

20 accepted papers

2026

OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding

CVPR 2026

We propose OpenVoxel, a training-free algorithm for grouping and captioning sparse voxels for the open-vocabulary 3D scene understanding tasks. Given the sparse voxel rasterization (SVR) model obtained from multi-view images of a 3D scene, our OpenVoxel is able to produce meaningful groups that desc

Cited by 0SourceScholar
2025

3D Gaussian Splatting with Grouped Uncertainty for Unconstrained Images

ICASSP 2025accepted

3D Gaussian Splatting (3DGS) [1] is a promising method for 3D reconstruction and novel view synthesis. However, training it with unconstrained images presents challenges due to transient objects that cause undesired floaters and ghosting artifacts. Although related works using Neural Radiance Fields…

Cited by 0SourceScholar
2025

BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL

ICLR 2025poster

Bayesian optimization (BO) offers an efficient pipeline for optimizing black-box functions with the help of a Gaussian process prior and an acquisition function (AF). Recently, in the context of single-objective BO, learning-based AFs witnessed promising empirical results given its favorable non-myo…

Cited by 0SourcePDFScholar
2025

FrugalNeRF: Fast Convergence for Extreme Few-shot Novel View Synthesis without Learned Priors

CVPR 2025poster

Neural Radiance Fields (NeRF) face significant challenges in extreme few-shot scenarios, primarily due to overfitting and long training times. Existing methods, such as FreeNeRF and SparseNeRF, use frequency regularization or pre-trained priors but struggle with complex scheduling and bias. We intro…

Cited by 0SourcePDFScholar
2025

LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos

ICCV 2025poster

LongSplat addresses critical challenges in novel view synthesis (NVS) from casually captured long videos characterized by irregular camera motion, unknown camera poses, and expansive scenes. Current methods often suffer from pose drift, inaccurate geometry initialization, and severe memory limitatio…

2025

Sparse Voxels Rasterization: Real-time High-fidelity Radiance Field Rendering

CVPR 2025poster

We propose an efficient radiance field rendering algorithm that incorporates a rasterization process on adaptive sparse voxels without neural networks or 3D Gaussians. There are two key contributions coupled with the proposed system. The first is to adaptively and explicitly allocate sparse voxels t…

2024

SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation

ECCV 2024poster

"We introduce SAM4MLLM, an innovative approach which integrates the Segment Anything Model (SAM) with Multi-Modal Large Language Models (MLLMs) for pixel-aware tasks. Our method enables MLLMs to learn pixel-level location information without requiring excessive modifications to the existing model ar…

2024

Seg2Reg: Differentiable 2D Segmentation to 1D Regression Rendering for 360 Room Layout Reconstruction

CVPR 2024poster

State-of-the-art single-view 360 room layout reconstruction methods formulate the problem as a high-level 1D (per-column) regression task. On the other hand traditional low-level 2D layout segmentation is simpler to learn and can represent occluded regions but it requires complex post-processing for…

2023

Hashing Neural Video Decomposition with Multiplicative Residuals in Space-Time

ICCV 2023poster

We present a video decomposition method that facilitates layer-based editing of videos with spatiotemporally varying lighting and motion effects. Our neural model decomposes an input video into multiple layered representations, each comprising a 2D texture map, a mask for the original video, and a…

Cited by 7PDFcodeScholar
2023

ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object Detection

ICCV 2023poster

We propose ImGeoNet, a multi-view image-based 3D object detection framework that models a 3D space by an image-induced geometry-aware voxel representation. Unlike previous methods which aggregate 2D features into 3D voxels without considering geometry, ImGeoNet learns to induce geometry from multi-…

Cited by 11PDFcodeScholar
2023

Neural-PBIR Reconstruction of Shape, Material, and Illumination

ICCV 2023poster

Reconstructing the shape and spatially varying surface appearances of a physical-world object as well as its surrounding illumination based on 2D images (e.g., photographs) of the object has been a long-standing problem in computer vision and graphics. In this paper, we introduce an accurate and hig…

Cited by 30PDFcodeScholar
2022

Direct Voxel Grid Optimization: Super-Fast Convergence for Radiance Fields Reconstruction

CVPR 2022oral

We present a super-fast convergence approach to reconstructing the per-scene radiance field from a set of images that capture the scene with known poses. This task, which is often applied to novel view synthesis, is recently revolutionized by Neural Radiance Field (NeRF) for its state-of-the-art qua…

Cited by 1232PDFcodeScholar
2021

Indoor Panorama Planar 3D Reconstruction via Divide and Conquer

CVPR 2021poster

Indoor panorama typically consists of human-made structures parallel or perpendicular to gravity. We leverage this phenomenon to approximate the scene in a 360-degree image with (H)orizontal-planes and (V)ertical-planes. To this end, we propose an effective divide-and-conquer strategy that divides p…

Cited by 16PDFcodeScholar
2021

Specialize and Fuse: Pyramidal Output Representation for Semantic Segmentation

ICCV 2021poster

We present a novel pyramidal output representation to ensure parsimony with our "specialize and fuse" process for semantic segmentation. A pyramidal "output" representation consists of coarse-to-fine levels, where each level is "specialize" in a different class distribution (e.g., more stuff than th…

Cited by 9PDFScholar
2019

HorizonNet: Learning Room Layout With 1D Representation and Pano Stretch Data Augmentation

CVPR 2019poster

We present a new approach to the problem of estimating the 3D room layout from a single panoramic image. We represent room layout as three 1D vectors that encode, at each image column, the boundary positions of floor-wall and ceiling-wall, and the existence of wall-wall boundary. The proposed networ…

Cited by 231PDFcodeScholar
2018

DLWV2: A Deep Learning-Based Wearable Vision-System with Vibrotactile-Feedback for Visually Impaired People to Reach Objects

IROS 2018poster

We develop a Deep Learning-based Wearable Vision-system with Vibrotactile-feedback (DLWV2)to guide Blind and Visually Impaired (BVI)people to reach objects. The system achieves high accuracy in object detection and tracking in 3-D using an extended deep learning-based 2.5-D detector and a 3-D object…

Cited by 19SourceScholar