TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classification, retrieval, segmentation and depth prediction. However, a fundamental capability that these models still struggle with is aligning dense patch r