RA-L 202414 citations

Self-Supervised Representation Learning From Temporal Ordering of Automated Driving Sequences

Christopher Lang, Alexander Braun, Lars Schillingmann, Karsten Haug, Abhinav Valada

Abstract

Self-supervised feature learning enables perception systems to benefit from the vast raw data recorded by vehicle fleets worldwide. While video-level self-supervised approaches have shown strong generalizability on classification tasks, the potential to learn dense representations from sequential data has been relatively unexplored. In this work, we propose <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">TempO</i> , a temporal ordering pretext task for pre-training region-level feature representations for perception tasks. We embed each frame by an unordered set of proposal feature vectors, a representation that is natural for object detection or tracking systems, and formulate the sequential ordering by predicting frame transition probabilities in a transformer-based multi-frame architecture whose complexity scales less than quadratic w.r.t. sequence length. Extensive evaluation in automated driving domains on the BDD100 K and MOT17 datasets shows that our <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">TempO</i> pre-training approach outperforms existing self-supervised single-frame methods as well as supervised transfer learning initialization strategies and achieves an improvement of <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$+0.7$</tex-math></inline-formula> in mAP for object detection and <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$+2.0\%$</tex-math></inline-formula> in the HOTA score for multi-object tracking.

BibTeX
@inproceedings{ral2024_selfsupervisedre,
  title = {Self-Supervised Representation Learning From Temporal Ordering of Automated Driving Sequences},
  author = {Christopher Lang and Alexander Braun and Lars Schillingmann and Karsten Haug and Abhinav Valada},
  booktitle = {RA-L 2024},
  year = {2024}
}
Self-Supervised Representation Learning From Temporal Ordering of Automated Driving Sequences · RA-L 2024