← Search

Marcello Federico

10 accepted papers

2025

MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation

ACL 2025long

Automatic evaluation of retrieval augmented generation (RAG) systems relies on fine-grained dimensions like faithfulness and relevance, as judged by expert human annotators. Meta-evaluation benchmarks support the development of automatic evaluators that correlate well with human judgement. However,…

2024

A Shocking Amount of the Web is Machine Translated: Insights from Multi-Way Parallelism

ACL 2024findings

We show that content on the web is often translated into many languages, and the low quality of these multi-way translations indicates they were likely created using Machine Translation (MT). Multi-way parallel, machine generated content not only dominates the translations in lower resource language…

2024

Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers’ Opinion Scores

ECCV 2024poster

"Recent advancements in audio-visual generative modeling have been propelled by progress in deep learning and the availability of data-rich benchmarks. However, the growth is not attributed solely to models and benchmarks. Universally accepted evaluation metrics also play an important role in advanc…

Cited by 1SourcePDFScholar
2023

End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation

EMNLP 2023long main

Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper, we tackle single-channel multi-speaker conversational ST with an end-to-end an…

Cited by 0SourcecodeScholar
2022

CoCoA-MT: A Dataset and Benchmark for Contrastive Controlled MT with Application to Formality

NAACL 2022findings

The machine translation (MT) task is typically formulated as that of returning a single translation for an input segment. However, in many cases, multiple different translations are valid and the appropriate translation may depend on the intended target audience, characteristics of the speaker, or e…

2022

Duration Modeling of Neural TTS for Automatic Dubbing

ICASSP 2022accepted

Automatic dubbing (AD) addresses the problem of translating speech in a video with speech in another language while preserving the viewer experience. A most important requirement of AD is isochrony, i.e. dubbed speech has to closely match the timing of speech and pauses of the original audio. In our…

Cited by 0SourceScholar
2022

ISOMETRIC MT: Neural Machine Translation for Automatic Dubbing

ICASSP 2022accepted

Automatic dubbing (AD) is among the machine translation (MT) use cases where translations should match a given length to allow for synchronicity between source and target speech. For neural MT, generating translations of length close to the source length (e.g. within ±10% in character count), while…

Cited by 0SourceScholar
2021

Improvements to Prosodic Alignment for Automatic Dubbing

ICASSP 2021accepted

Automatic dubbing is an extension of speech-to-speech translation such that the resulting target speech is carefully aligned in terms of duration, lip movements, timbre, emotion, prosody, etc. of the speaker in order to achieve audiovisual coherence. Dubbing quality strongly depends on isochrony, i.…

Cited by 0SourceScholar
2021

Machine Translation Verbosity Control for Automatic Dubbing

ICASSP 2021accepted

Automatic dubbing aims at seamlessly replacing the speech in a video document with synthetic speech in a different language. The task implies many challenges, one of which is generating translations that not only convey the original content, but also match the duration of the corresponding utterance…

Cited by 0SourceScholar