Enhancing Vision Transformers for Object Detection via Context-Aware Token Selection and Packing
In recent years, the long-range attention mechanism of vision transformers has driven significant performance breakthroughs across various computer vision tasks. However, these advancements come at the cost of inefficiency and substantial computational expense, especially when dealing with sparse da…