2026
SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation
CVPR 2026
Foundational Vision Transformers (ViTs) have limited effectiveness in tasks requiring fine-grained spatial understanding, due to their fixed pre-training resolution and inherently coarse patch-level representations. These challenges are especially pronounced in dense prediction scenarios, such as op