← Search

Lincan Cai

4 accepted papers

2025

From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection

ICML 2025poster

Pretrained vision-language models (VLMs), e.g., CLIP, demonstrate impressive zero-shot capabilities on downstream tasks. Prior research highlights the crucial role of visual augmentation techniques, like random cropping, in alignment with fine-grained class descriptions generated by large language m…

2024

Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation

ICML 2024poster

Large-scale pretrained models have proven immensely valuable in handling data-intensive modalities like text and image. However, fine-tuning these models for certain specialized modalities, such as protein sequence and cosmic ray, poses challenges due to the significant modality discrepancy and scar…

Cited by 4SourcePDFScholar
2024

Learning Modality Knowledge Alignment for Cross-Modality Transfer

ICML 2024poster

Cross-modality transfer aims to leverage large pretrained models to complete tasks that may not belong to the modality of pretraining data. Existing works achieve certain success in extending classical finetuning to cross-modal scenarios, yet we still lack understanding about the influence of modali…

Cited by 3SourcePDFScholar
2023

Language Semantic Graph Guided Data-Efficient Learning

NeurIPS 2023poster

Developing generalizable models that can effectively learn from limited data and with minimal reliance on human supervision is a significant objective within the machine learning community, particularly in the era of deep neural networks. Therefore, to achieve data-efficient learning, researchers ty…