2024
HaltingVT: Adaptive Token Halting Transformer for Efficient Video Recognition
ICASSP 2024accepted
Action recognition in videos poses a challenge due to its high computational cost, especially for Joint Space-Time video transformers (Joint VT). Despite their effectiveness, the excessive number of tokens in such architectures significantly limits their efficiency. In this paper, we propose Halting…