2025
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
NeurIPS 2025poster
Vision transformers are ever larger, more accurate, and more expensive to compute. At high resolution, the expense is even more extreme as the number of tokens grows quadratically in the image size. We turn to adaptive computation to cope with this cost by learning to predict where to compute. Our…