← Search

Linghao Jin

5 accepted papers

2026

Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination

ICML 2026oral

Vision-Language Models (VLMs) frequently generate self-reflective statements during reasoning, such as ``let me check the figure again.'' Do such statements trigger genuine visual re-examination, or merely represent learned textual patterns? We investigate this question through VisualSwap, an image-…

Cited by 3SourceScholar
2025

Progressive Compositionality in Text-to-Image Generative Models

ICLR 2025spotlight

Despite the impressive text-to-image (T2I) synthesis capabilities of diffusion models, they often struggle to understand compositional relationships between objects and attributes, especially in complex settings. Existing approaches through building compositional architectures or generating difficul…

2024

Light-weight Fine-tuning Method for Defending Adversarial Noise in Pre-trained Medical Vision-Language Models

EMNLP 2024finding

Fine-tuning pre-trained Vision-Language Models (VLMs) has shown remarkable capabilities in medical image and textual depiction synergy. Nevertheless, many pre-training datasets are restricted by patient privacy concerns, potentially containing noise that can adversely affect downstream performance.…

Cited by 2SourcePDFScholar
2021

Domain Generalization under Conditional and Label Shifts via Variational Bayesian Inference

IJCAI 2021poster

In this work, we propose a domain generalization (DG) approach to learn on several labeled source domains and transfer knowledge to a target domain that is inaccessible in training. Considering the inherent conditional and label shifts, we would expect the alignment of p(x|y) and p(y). However, the…

Cited by 33SourcePDFScholar