← Search

Selen Pehlivan

2 accepted papers

2025

Learning to Describe Implicit Changes: Noise-robust Pre-training for Image Difference Captioning

EMNLP 2025

Image Difference Captioning (IDC) methods have advanced in highlighting subtle differences between similar images, but their performance is often constrained by limited training data. Using Large Multimodal Models (LMMs) to describe changes in image pairs mitigates data limits but adds noise. These

Cited by 0SourcePDFScholar
2024

Text-to-Multimodal Retrieval with Bimodal Input Fusion in Shared Cross-Modal Transformer

COLING 2024main

The rapid proliferation of multimedia content has necessitated the development of effective multimodal video retrieval systems. Multimodal video retrieval is a non-trivial task involving retrieval of relevant information across different modalities, such as text, audio, and visual. This work aims to…