IROS 2023poster5 citations

Scaling Vision-Based End-to-End Autonomous Driving with Multi-View Attention Learning

Yi Xiao, Felipe Codevilla, Diego Porres, Antonio M. López

Abstract

On end-to-end driving, human driving demonstrations are used to train perception-based driving models by imitation learning. This process is supervised on vehicle signals (e.g., steering angle, acceleration) but does not require extra costly supervision (human labeling of sensor data). As a representative of such vision-based end-to-end driving models, CILRS is commonly used as a baseline to compare with new driving models. So far, some latest models achieve better performance than CILRS by using expensive sensor suites and/or by using large amounts of human-labeled data for training. Given the difference in performance, one may think that it is not worth pursuing vision-based pure end-to-end driving. However, we argue that this approach still has great value and potential considering cost and maintenance. In this paper, we present CIL++, which improves on CILRS by both processing higher-resolution images using a human-inspired HFOV as an inductive bias and incorporating a proper attention mechanism. CIL++ achieves competitive performance compared to models which are more costly to develop. We propose to replace CILRS with CIL++ as a strong vision-based pure end-to-end driving baseline supervised by only vehicle signals and trained by conditional imitation learning.

BibTeX
@inproceedings{iros2023_scalingvisionbas,
  title = {Scaling Vision-Based End-to-End Autonomous Driving with Multi-View Attention Learning},
  author = {Yi Xiao and Felipe Codevilla and Diego Porres and Antonio M. López},
  booktitle = {IROS 2023},
  year = {2023}
}
Scaling Vision-Based End-to-End Autonomous Driving with Multi-View Attention Learning · IROS 2023