← Search

Arnas Uselis

8 accepted papers

2026

CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally

ICLR 2026poster

CLIP (Contrastive Language-Image Pretraining) has become a popular choice for various downstream tasks. However, recent studies have questioned its ability to represent compositional concepts effectively. These works suggest that CLIP often acts like a bag-of-words (BoW) model, interpreting images a…

Cited by 0SourcecodeScholar
2026

When Do Diffusion Models learn to Generate Multiple Objects?

ICML 2026poster

Text-to-image diffusion models achieve impressive visual fidelity, yet they remain unreliable in multi-object generation. Despite extensive empirical evidence of these failures, the underlying causes remain unclear. We begin by asking how much of this limitation arises from the data itself. To disen…

Cited by 0SourceScholar
2025

Diffusion Classifiers Understand Compositionality, but Conditions Apply

NeurIPS 2025poster

Understanding visual scenes is fundamental to human intelligence. While discriminative models have significantly advanced computer vision, they often struggle with compositional understanding. In contrast, recent generative text-to-image diffusion models excel at synthesizing complex scenes, suggest…

Cited by 0SourcecodeScholar