Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
Current state-of-the-art segmentation models encode entire images before focusing on specific objects. This wastes computational resources. We introduce FLIP (Fovea-Like Input Patching), a parameter-efficient vision model that realizes object segmentation through biologically-inspired top-down atten…