← Search

Sohwi Lim

3 accepted papers

2026

CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space

CVPR 2026

Human perception of visual similarity is inherently adaptive and subjective, depending on the users' interests and focus. However, most image retrieval systems fail to reflect this flexibility, relying on a fixed, monolithic metric that cannot incorporate multiple conditions simultaneously. To addre

Cited by 0SourceScholar
2026

Measurement-Consistent Langevin Corrector for Stabilizing Latent Diffusion Inverse Problem Solvers

ICML 2026poster

While latent diffusion models (LDMs) have emerged as powerful priors for inverse problems, existing LDM-based solvers frequently suffer from instability. In this work, we first identify the instability as a discrepancy between the solver dynamics and stable reverse diffusion dynamics learned by the …

Cited by 1SourceScholar
2026

Zero-Shot Rankability: Revealing Latent Ordinal Structure in Multimodal Large Language Models via Language

ICML 2026poster

Recent work shows that vision encoders capture ordinal attributes along linear axes, which can be recovered from as few as two labeled images. However, in the zero-shot setting, the text-driven rank axis for Vision-Language Models (VLMs) like CLIP remains suboptimal. In this work, we study the embed…

Cited by 0SourceScholar