← Search

Kevin Duh

20 accepted papers

2026

Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG

ICML 2026spotlight

Multilingual Retrieval-Augmented Generation (mRAG) systems enable language models to answer knowledge-intensive queries with citation-supported responses across languages. Despite their growing use, an open questions is whether the mixture of different document languages impacts generation and citat…

Cited by 0SourceScholar
2025

An Interdisciplinary Approach to Human-Centered Machine Translation

EMNLP 2025

Machine Translation (MT) tools are widely used today, often in contexts where professional translators are not present. Despite progress in MT technology, a gap persists between system development and real-world usage, particularly for non-expert users who may struggle to assess translation reliabil

Cited by 0SourcePDFScholar
2025

HLTCOE Submission to the VoicePrivacy Attacker Challenge

ICASSP 2025accepted

We describe our submission to the 2024 VoicePrivacy Attacker Challenge. We propose three main categories of methods to improve ASV performance against anonymized speech: improvements to the underlying classifier, alternative distance metrics when computing ASV scores, and kNN-VC normalization. By si…

Cited by 0SourceScholar
2025

Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation

EMNLP 2025

As people increasingly use AI systems in work and daily life, feedback mechanisms that help them use AI responsibly are urgently needed, particularly in settings where users are not equipped to assess the quality of AI predictions. We study a realistic Machine Translation (MT) scenario where monolin

2025

Whisper-UT: A Unified Translation Framework for Speech and Text

EMNLP 2025

Encoder-decoder models have achieved remarkable success in speech and text tasks, yet efficiently adapting these models to diverse uni/multi-modal scenarios remains an open challenge. In this paper, we propose Whisper-UT, a unified and efficient framework that leverages lightweight adapters to enabl

Cited by 0SourcePDFScholar
2024

Anti-LM Decoding for Zero-shot In-context Machine Translation

NAACL 2024findings

Zero-shot In-context learning is the phenomenon where models can perform a task given only the instructions. However, pre-trained large language models are known to be poorly calibrated for zero-shot tasks. One of the most effective approaches to handling this bias is to adopt a contrastive decoding…

2024

Exploring Geometric Representational Disparities between Multilingual and Bilingual Translation Models

COLING 2024main

Multilingual machine translation has proven immensely useful for both parameter efficiency and overall performance across many language pairs via complete multilingual parameter sharing. However, some language pairs in multilingual models can see worse performance than in bilingual models, especiall…

Cited by 0SourcePDFScholar
2023

Handshape-Aware Sign Language Recognition: Extended Datasets and Exploration of Handshape-Inclusive Methods

EMNLP 2023long findings

The majority of existing work on sign language recognition encodes signed videos without explicitly acknowledging the phonological attributes of signs. Given that handshape is a vital parameter in sign languages, we explore the potential of handshape-aware sign language recognition. We augment the P…

Cited by 0SourceScholar
2022

AfriCLIRMatrix: Enabling Cross-Lingual Information Retrieval for African Languages

EMNLP 2022main

Language diversity in NLP is critical in enabling the development of tools for a wide range of users.However, there are limited resources for building such tools for many languages, particularly those spoken in Africa.For search, most existing datasets feature few or no African languages, directly i…

2022

Bilingual Lexicon Induction for Low-Resource Languages using Graph Matching via Optimal Transport

EMNLP 2022main

Bilingual lexicons form a critical component of various natural language processing applications, including unsupervised and semisupervised machine translation and crosslingual information retrieval. In this work, we improve bilingual lexicon induction performance across 40 language pairs with a gra…

2022

IsoVec: Controlling the Relative Isomorphism of Word Embedding Spaces

EMNLP 2022main

The ability to extract high-quality translation dictionaries from monolingual word embedding spaces depends critically on the geometric similarity of the spaces—their degree of “isomorphism.” We address the root-cause of faulty cross-lingual mapping: that word embedding training resulted in the unde…

2022

Offer a Different Perspective: Modeling the Belief Alignment of Arguments in Multi-party Debates

EMNLP 2022main

In contexts where debate and deliberation are the norm, the participants are regularly presented with new information that conflicts with their original beliefs. When required to update their beliefs (belief alignment), they may choose arguments that align with their worldview (confirmation bias). W…

2021

An Analysis of Euclidean vs. Graph-Based Framing for Bilingual Lexicon Induction from Word Embedding Spaces

EMNLP 2021finding

Much recent work in bilingual lexicon induction (BLI) views word embeddings as vectors in Euclidean space. As such, BLI is typically solved by finding a linear transformation that maps embeddings to a common space. Alternatively, word embeddings may be understood as nodes in a weighted graph. This f…

2021

Data and Parameter Scaling Laws for Neural Machine Translation

EMNLP 2021main

We observe that the development cross-entropy loss of supervised neural machine translation models scales like a power law with the amount of training data and the number of non-embedding parameters in the model. We discuss some practical implications of these results, such as predicting BLEU achiev…

2021

ORTHROS: non-autoregressive end-to-end speech translation With dual-decoder

ICASSP 2021accepted

Fast inference speed is an important goal towards real-world deployment of speech translation (ST) systems. End-to-end (E2E) models based on the encoder-decoder architecture are more suitable for this goal than traditional cascaded systems, but their effectiveness regarding decoding speed has not be…

Cited by 0SourceScholar
2018

Audio-Visual Person Recognition in Multimedia Data From the Iarpa Janus Program

ICASSP 2018accepted

Currently, datasets that support audio-visual recognition of people in videos are scarce and limited. In this paper, we introduce an expansion of video data from the IARPA Janus program to support this research area. We refer to the expanded set, which adds labels for voice to the already-existing f…

Cited by 0SourceScholar