← Search

Lingjie Yi

3 accepted papers

2026

The Convergent Representation of Vision-Language Contrastive Learning: Geometry, Modality Gap and Shared Space Alignment

ICML 2026poster

Multimodal contrastive learning (MCL) aims to embed data from two modalities in a shared embedding space. However, in practice, representations of images and text occupy completely separate regions of embedding space, a phenomenon called the modality gap. Moreover, experimental findings on how the s…

Cited by 0SourceScholar
2025

Backdooring Vision-Language Models with Out-Of-Distribution Data

ICLR 2025poster

The emergence of Vision-Language Models (VLMs) represents a significant advancement in integrating computer vision with Large Language Models (LLMs) to generate detailed text descriptions from visual inputs. Despite their growing importance, the security of VLMs, particularly against backdoor attack…

Cited by 3SourcePDFScholar
2025

Geometry of Long-Tailed Representation Learning: Rebalancing Features for Skewed Distributions

ICLR 2025poster

Deep learning has achieved significant success by training on balanced datasets. However, real-world data often exhibit long-tailed distributions. Empirical studies have revealed that long-tailed data skew data representations, where head classes dominate the feature space. Many methods have been pr…

Cited by 0SourcePDFScholar