← Search

David Etter

3 accepted papers

2025

MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval

CVPR 2025poster

Efficiently retrieving and synthesizing information from large-scale multimodal collections has become a critical challenge. However, existing video retrieval datasets suffer from scope limitations, primarily focusing on matching descriptive but vague queries with small collections of professionally…

Cited by 1SourcePDFScholar
2024

Grounding Partially-Defined Events in Multimodal Data

EMNLP 2024finding

How are we able to learn about complex current events just from short snippets of video? While natural language enables straightforward ways to represent under-specified, partially observable events, visual data does not facilitate analogous methods and, consequently, introduces unique challenges in…

Cited by 1SourcePDFScholar
2021

Robust Open-Vocabulary Translation from Visual Text Representations

EMNLP 2021main

Machine translation models have discrete vocabularies and commonly use subword segmentation techniques to achieve an ‘open vocabulary.’ This approach relies on consistent and correct underlying unicode sequences, and makes models susceptible to degradation from common types of noise and variation. M…