← Search

Avijit Vajpayee

2 accepted papers

2025

Detect, Disambiguate, and Translate: On-Demand Visual Reasoning for Multimodal Machine Translation with Large Vision-Language Models

NAACL 2025long

Multimodal machine translation (MMT) aims to leverage additional modalities to assist in language translation. With limited parallel data, current MMT systems rely heavily on monolingual English captioning data. These systems face three key issues: they often overlook that visual signals are unneces…

Cited by 0SourcePDFScholar
2023

MEGA: Multimodal Alignment Aggregation and Distillation For Cinematic Video Segmentation

ICCV 2023poster

Previous research has studied the task of segmenting cinematic videos into scenes and into narrative acts. However, these studies have overlooked the essential task of multimodal alignment and fusion for effectively and efficiently processing long-form videos (>60min). In this paper, we introduce Mu…

Cited by 4PDFcodeScholar