← Search

Chunyu Xie

6 accepted papers

2026

AMLRIS: Alignment-aware Masked Learning for Referring Image Segmentation

ICLR 2026poster

Referring Image Segmentation (RIS) aims to segment the object in an image uniquely referred to by a natural language expression. However, RIS training often contains hard-to-align and instance-specific visual signals; optimizing on such pixels injects misleading gradients and drives the model in the…

Cited by 0SourcecodeScholar
2026

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

ICML 2026poster

Fine-grained vision-language understanding requires precise alignment between visual content and linguistic descriptions, a capability that remains limited in current models, particularly in non-English settings. While models like CLIP perform well on global alignment, they often struggle to capture…

Cited by 0SourceScholar
2025

FG-CLIP: Fine-Grained Visual and Textual Alignment

ICML 2025poster

Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short captions. To address this, we propose Fine-Grained CLIP (FG-CLIP), which enhances…

2025

IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities

AAAI 2025technical

In the field of multimodal large language models (MLLMs), common methods typically involve unfreezing the language model during training to foster profound visual understanding. However, the fine-tuning of such models with vision-language data often leads to a diminution of their natural language pr…

2025

LMM-Det: Make Large Multimodal Models Excel in Object Detection

ICCV 2025poster

Large multimodal models (LMMs) have garnered wide-spread attention and interest within the artificial intelligence research and industrial communities, owing to their remarkable capability in multimodal understanding, reasoning, and in-context learning, among others. While LMMs have demonstrated pro…

2025

Prompt as Knowledge Bank: Boost Vision-language model via Structural Representation for zero-shot medical detection

ICLR 2025poster

Zero-shot medical detection can further improve detection performance without relying on annotated medical images even upon the fine-tuned model, showing great clinical value. Recent studies leverage grounded vision-language models (GLIP) to achieve this by using detailed disease descriptions as pro…

Cited by 0SourcePDFScholar