ICRA 2026poster0 citations

PAGTM: Position and Attention-Guided Token Merging for Efficient Visual Place Recognition

Hongchan Cho, Youngjo Lee, Jinwoo Jang, Seunghan Yu, Euntai Kim

Abstract

Recent advances in Vision Transformers (ViTs) have significantly improved the performance of Visual Place Recognition (VPR), but their high computational cost—due to the quadratic complexity of self-attention—limits their practical deployment in real-world scenarios. To address this challenge, we propose PAGTM (Positional- and Attention-Guided Token Merging), a training-free token reduction framework designed specifically for ViT-based VPR models. In VPR, preserving the spatial layout of a scene (e.g. road alignment, building structures) and focusing on semantically meaningful regions are both critical for robust matching under viewpoint and appearance variations. However, existing token reduction methods often overlook these aspects, leading to degraded recognition performance. To address this, PAGTM incorporates two key cues. The first is positional proximity, which merges spatially adjacent tokens to maintain the scene’s structural layout. The second is attention-based token protection, which retains tokens that receive high attention because they represent regions important for distinguishing places, such as signs or distinctive structures. Without requiring any fine-tuning, PAGTM can be directly applied at inference time and consistently outperforms existing token reduction methods such as ToMe and ToFu across multiple ViT-based VPR models and datasets, achieving a better trade-off between computational efficiency and retrieval accuracy.

Deep Learning for Visual PerceptionRecognitionLocalization
PAGTM: Position and Attention-Guided Token Merging for Efficient Visual Place Recognition · ICRA 2026