← Search

Ji Ao

2 accepted papers

2026

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

ICML 2026poster

Fine-grained vision-language understanding requires precise alignment between visual content and linguistic descriptions, a capability that remains limited in current models, particularly in non-English settings. While models like CLIP perform well on global alignment, they often struggle to capture…

Cited by 0SourceScholar
2025

LMM-Det: Make Large Multimodal Models Excel in Object Detection

ICCV 2025poster

Large multimodal models (LMMs) have garnered wide-spread attention and interest within the artificial intelligence research and industrial communities, owing to their remarkable capability in multimodal understanding, reasoning, and in-context learning, among others. While LMMs have demonstrated pro…