← Search

Sheng-Yu Huang

6 accepted papers

2026

OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding

CVPR 2026

We propose OpenVoxel, a training-free algorithm for grouping and captioning sparse voxels for the open-vocabulary 3D scene understanding tasks. Given the sparse voxel rasterization (SVR) model obtained from multi-view images of a 3D scene, our OpenVoxel is able to produce meaningful groups that desc

Cited by 0SourceScholar
2025

3D Gaussian Inpainting with Depth-Guided Cross-View Consistency

CVPR 2025poster

When performing 3D inpainting using novel-view rendering methods like Neural Radiance Field (NeRF) or 3D Gussian Splatting (3DGS), how to achieve texture and geometry consistency across camera views has been a challenge. In this paper, we propose a framework of 3D Gaussian Inpainting with Depth-Guid…

Cited by 0SourcePDFScholar
2024

Developing the Keep-Important-Samples Scheme for Training the Advanced CNN-Based Automatic Virtual Metrology Models

RA-L 2024

Virtual Metrology (VM) technology can convert offline sampling inspection into online and real-time total inspection. As the processes of high-tech industries (semiconductor or TFT-LCD) are getting more sophisticated, higher VM prediction accuracy is demanded. With regard to this requirement, the ad

Cited by 5SourceScholar
2024

GSNeRF: Generalizable Semantic Neural Radiance Fields with Enhanced 3D Scene Understanding

CVPR 2024poster

Utilizing multi-view inputs to synthesize novel-view images Neural Radiance Fields (NeRF) have emerged as a popular research topic in 3D vision. In this work we introduce a Generalizable Semantic Neural Radiance Field (GSNeRF) which uniquely takes image semantics into the synthesis process so that b…

Cited by 8SourcePDFScholar
2024

TPA3D: Triplane Attention for Fast Text-to-3D Generation

ECCV 2024poster

"Due to the lack of large-scale text-3D correspondence data, recent text-to-3D generation works mainly rely on utilizing 2D diffusion models for synthesizing 3D data. Since diffusion-based methods typically require significant optimization time for both training and inference, the use of GAN-based m…

Cited by 3SourcePDFScholar
2020

Convolution in the Cloud: Learning Deformable Kernels in 3D Graph Convolution Networks for Point Cloud Analysis

CVPR 2020poster

Point clouds are among the popular geometry representations for 3D vision applications. However, without regular structures like 2D images, processing and summarizing information over these unordered data points are very challenging. Although a number of previous works attempt to analyze point cloud…

Cited by 277PDFcodeScholar