← Search

Simon Gottschalk

2 accepted papers

2025

Aligning Visual Contrastive learning models via Preference Optimization

ICLR 2025poster

Contrastive learning models have demonstrated impressive abilities to capture semantic similarities by aligning representations in the embedding space. However, their performance can be limited by the quality of the training data and its inherent biases. While Preference Optimization (PO) methods su…

2025

Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders

NeurIPS 2025poster

Despite significant advances in Multimodal Large Language Models (MLLMs), understanding complex temporal dynamics in videos remains a major challenge. Our experiments show that current Video Large Language Model (Video-LLM) architectures have critical limitations in temporal understanding, strugglin…

Cited by 0SourcecodeScholar