ICRA 2024poster0 citations

CVFormer: Learning Circum-View Representation and Consistency for Vision-Based Occupancy Prediction via Transformers

Zhengqi Bai, Wenjun Shi, Dongchen Zhu, Hanlong Kang, Guanghui Zhang, Gang Ye, Yang Xiao, Lei Wang

Abstract

With the increasing demands for perception accuracy in autonomous driving, there is a growing focus on fine-grained 3D semantic occupancy prediction. Effectively representing detailed three-dimensional scenes has become a significant challenge in the development of this task. In this paper, we present a novel transformer-based framework named CVFormer, which leverages two-dimensional circum-views from the ego to excavate three-dimensional features of the surrounding environment. Circum-views provide a novel solution for effectively addressing the representation of dense and fine-grained scenes. Specifically, a multi-attention module CTMA is designed for fusing temporal features from circum-views to fully exploit the spatiotemporal correlations between frames and capture more comprehensive clues. Furthermore, a novel 2D projection constraint is established by observing objects from different perspective directions, and multiple 3D constraints based on object invariance and semantic consistency are also conducted for supervising the network, which enhances its performance of understanding the scene. Experimental results on nuScenes dataset demonstrate that the proposed CVFormer obviously outperforms existing methods for occupancy prediction.

BibTeX
@inproceedings{icra2024_cvformerlearning,
  title = {CVFormer: Learning Circum-View Representation and Consistency for Vision-Based Occupancy Prediction via Transformers},
  author = {Zhengqi Bai and Wenjun Shi and Dongchen Zhu and Hanlong Kang and Guanghui Zhang and Gang Ye and Yang Xiao and Lei Wang and Xiaolin Zhang and Bo Li and Jiamao Li},
  booktitle = {ICRA 2024},
  year = {2024}
}
CVFormer: Learning Circum-View Representation and Consistency for Vision-Based Occupancy Prediction via Transformers · ICRA 2024