← Search

Reno Kriz

7 accepted papers

2025

MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval

CVPR 2025poster

Efficiently retrieving and synthesizing information from large-scale multimodal collections has become a critical challenge. However, existing video retrieval datasets suffer from scope limitations, primarily focusing on matching descriptive but vague queries with small collections of professionally…

Cited by 1SourcePDFScholar
2025

Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval

CVPR 2025poster

In this work, we tackle the problem of text-to-video retrieval (T2VR). Inspired by the success of late interaction techniques in text-document, text-image, and text-video retrieval, our approach, Video-ColBERT, introduces a simple and efficient mechanism for fine-grained similarity assessment betwee…

Cited by 0SourcePDFScholar
2025

Whisper-UT: A Unified Translation Framework for Speech and Text

EMNLP 2025

Encoder-decoder models have achieved remarkable success in speech and text tasks, yet efficiently adapting these models to diverse uni/multi-modal scenarios remains an open challenge. In this paper, we propose Whisper-UT, a unified and efficient framework that leverages lightweight adapters to enabl

Cited by 0SourcePDFScholar
2024

Grounding Partially-Defined Events in Multimodal Data

EMNLP 2024finding

How are we able to learn about complex current events just from short snippets of video? While natural language enables straightforward ways to represent under-specified, partially observable events, visual data does not facilitate analogous methods and, consequently, introduces unique challenges in…

Cited by 1SourcePDFScholar
2023

On Event Individuation for Document-Level Information Extraction

EMNLP 2023short findings

As information extraction (IE) systems have grown more adept at processing whole documents, the classic task of *template filling* has seen renewed interest as a benchmark for document-level IE. In this position paper, we call into question the suitability of template filling for this purpose. We ar…

Cited by 0SourcecodeScholar
2022

Ambiguous Images With Human Judgments for Robust Visual Event Classification

NeurIPS 2022accept

Contemporary vision benchmarks predominantly consider tasks on which humans can achieve near-perfect performance. However, humans are frequently presented with visual data that they cannot classify with 100% certainty, and models trained on standard vision benchmarks achieve low performance when eva…

Cited by 12SourcePDFScholar
2021

BiSECT: Learning to Split and Rephrase Sentences with Bitexts

EMNLP 2021main

An important task in NLP applications such as sentence simplification is the ability to take a long, complex sentence and split it into shorter sentences, rephrasing as necessary. We introduce a novel dataset and a new model for this ‘split and rephrase’ task. Our BiSECT training data consists of 1…