← Search

Ivona Najdenkoska

2 accepted papers

2025

TULIP: Token-length Upgraded CLIP

ICLR 2025poster

We address the challenge of representing long captions in vision-language models, such as CLIP. By design these models are limited by fixed, absolute positional encodings, restricting inputs to a maximum of 77 tokens and hindering performance on tasks requiring longer descriptions. Although recent w…

2023

Meta Learning to Bridge Vision and Language Models for Multimodal Few-Shot Learning

ICLR 2023poster

Multimodal few-shot learning is challenging due to the large domain gap between vision and language modalities. Existing methods are trying to communicate visual concepts as prompts to frozen language models, but rely on hand-engineered task induction to reduce the hypothesis space. To make the whol…