2025
TIPS: Text-Image Pretraining with Spatial awareness
ICLR 2025poster
While image-text representation learning has become very popular in recent years, existing models tend to lack spatial awareness and have limited direct applicability for dense understanding tasks. For this reason, self-supervised image-only pretraining is still the go-to method for many dense visio…