PAGTM: Position and Attention-Guided Token Merging for Efficient Visual Place Recognition
Recent advances in Vision Transformers (ViTs) have significantly improved the performance of Visual Place Recognition (VPR), but their high computational cost—due to the quadratic complexity of self-attention—limits their practical deployment in real-world scenarios. To address this challenge, we pr…