BEVoxSeg: BEV-Voxel Representation for Fast and Accurate Camera-Based 3D Segmentation
Haiyi Liu, Beibei Wang, Lu Zhang, Jianmin Ji, Yanyong Zhang
Abstract
Recent research has demonstrated the advantages of Bird’s-eye-view (BEV) representation in the field of 3D perception. However, due to the lack of height information, BEV representation alone is insufficient to accurately reconstruct the complete surrounding 3D scene. On the other hand, voxel representation excels in describing 3D structures, but their memory and computational cost pose challenges for fast inference. To tackle these limitations, we propose an innovative method dubbed BEVoxSeg, which leverages the computational efficiency of BEV methods while incorporating essential geometric information from voxel features. By combining the advantages from both representations, our approach achieved state-of-the-art results for LiDAR semantic segmentation on nuScenes and demonstrated a superior performance in the occupancy prediction tasks on Occ3D-nuScenes dataset.
BibTeX
@inproceedings{icassp2024_bevoxsegbevvoxel,
title = {BEVoxSeg: BEV-Voxel Representation for Fast and Accurate Camera-Based 3D Segmentation},
author = {Haiyi Liu and Beibei Wang and Lu Zhang and Jianmin Ji and Yanyong Zhang},
booktitle = {ICASSP 2024},
year = {2024}
}