Flexible and Efficient Spatio-Temporal Transformer for Sequential Visual Place Recognition
Sequential Visual Place Recognition (Seq-VPR) leverages transformers to capture spatio-temporal features effectively. In practice, a transformer-based Seq-VPR model should be flexible to the number of frames per sequence (sequence length), deliver fast inference, and use little memory to meet real-t…