← Search

Fanjing Kong

2 accepted papers

2026

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

ICML 2026poster

Fine-grained vision-language understanding requires precise alignment between visual content and linguistic descriptions, a capability that remains limited in current models, particularly in non-English settings. While models like CLIP perform well on global alignment, they often struggle to capture…

Cited by 0SourceScholar
2025

FG-CLIP: Fine-Grained Visual and Textual Alignment

ICML 2025poster

Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short captions. To address this, we propose Fine-Grained CLIP (FG-CLIP), which enhances…