AAAI 2024technical6 citations

Self-Supervised Bird’s Eye View Motion Prediction with Cross-Modality Signals

Shaoheng Fang, Zuhong Liu, Mingyu Wang, Chenxin Xu, Yiqi Zhong, Siheng Chen

Abstract

Learning the dense bird's eye view (BEV) motion flow in a self-supervised manner is an emerging research for robotics and autonomous driving. Current self-supervised methods mainly rely on point correspondences between point clouds, which may introduce the problems of fake flow and inconsistency, hindering the model’s ability to learn accurate and realistic motion. In this paper, we introduce a novel cross-modality self-supervised training framework that effectively addresses these issues by leveraging multi-modality data to obtain supervision signals. We design three innovative supervision signals to preserve the inherent properties of scene motion, including the masked Chamfer distance loss, the piecewise rigidity loss, and the temporal consistency loss. Through extensive experiments, we demonstrate that our proposed self-supervised framework outperforms all previous self-supervision methods for the motion prediction task.

BibTeX
@article{Fang_Liu_Wang_Xu_Zhong_Chen_2024, title={Self-Supervised Bird’s Eye View Motion Prediction with Cross-Modality Signals}, volume={38}, url={https://ojs.aaai.org/index.php/AAAI/article/view/27940}, DOI={10.1609/aaai.v38i2.27940}, abstractNote={Learning the dense bird’s eye view (BEV) motion flow in a self-supervised manner is an emerging research for robotics and autonomous driving. Current self-supervised methods mainly rely on point correspondences between point clouds, which may introduce the problems of fake flow and inconsistency, hindering the model’s ability to learn accurate and realistic motion. In this paper, we introduce a novel cross-modality self-supervised training framework that effectively addresses these issues by leveraging multi-modality data to obtain supervision signals. We design three innovative supervision signals to preserve the inherent properties of scene motion, including the masked Chamfer distance loss, the piecewise rigidity loss, and the temporal consistency loss. Through extensive experiments, we demonstrate that our proposed self-supervised framework outperforms all previous self-supervision methods for the motion prediction task.}, number={2}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Fang, Shaoheng and Liu, Zuhong and Wang, Mingyu and Xu, Chenxin and Zhong, Yiqi and Chen, Siheng}, year={2024}, month={Mar.}, pages={1726-1734} }
Self-Supervised Bird’s Eye View Motion Prediction with Cross-Modality Signals · AAAI 2024