← Search

Moran Yanuka

5 accepted papers

2025

Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions

NAACL 2025long

Recent research increasingly focuses on training vision-language models (VLMs) with long, detailed image captions. However, small-scale VLMs often struggle to balance the richness of these captions with the risk of hallucinating content during fine-tuning. In this paper, we explore how well VLMs ada…

2025

EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits

ACL 2025long

Text-guided image editing, fueled by recent advancements in generative AI, is becoming increasingly widespread. This trend highlights the need for a comprehensive framework to verify text-guided edits and assess their quality. To address this need, we introduce EditInspector, a novel benchmark for e…

Cited by 0SourcePDFScholar
2024

ICC : Quantifying Image Caption Concreteness for Multimodal Dataset Curation

ACL 2024findings

Web-scale training on paired text-image data is becoming increasingly central to multimodal learning, but is challenged by the highly noisy nature of datasets in the wild. Standard data filtering approaches succeed in removing mismatched text-image pairs, but permit semantically related but highly a…

2024

Mitigating Open-Vocabulary Caption Hallucinations

EMNLP 2024main

While recent years have seen rapid progress in image-conditioned text generation, image captioning still suffers from the fundamental issue of hallucinations, namely, the generation of spurious details that cannot be inferred from the given image. Existing methods largely use closed-vocabulary objec…