LaCViT: A Label-Aware Contrastive Fine-Tuning Framework for Vision Transformers
Vision Transformers (ViTs) have emerged as popular models in computer vision, demonstrating state-of-the-art performance across various tasks. This success typically follows a two-stage strategy involving pre-training on large-scale datasets using self-supervised signals, such as masked random patch…