← Search

Dietrich Klakow

41 accepted papers

2026

Split Personality Training: Revealing Latent Knowledge Through Alternate Personalities

ICML 2026poster

Detecting misalignment in large language models is challenging because models may learn to conceal misbehavior during training. Standard auditing techniques fall short: black-box methods often cannot distinguish misaligned outputs from benign ones, and mechanistic interpretability does not scale wit…

Cited by 0SourceScholar
2025

AFRIDOC-MT: Document-level MT Corpus for African Languages

EMNLP 2025

This paper introduces AFRIDOC-MT, a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yorùbá, and Zulu. The dataset comprises 334 health and 271 information technology news documents, all human-translated from English to these lan

2025

Accept or Deny? Evaluating LLM Fairness and Performance in Loan Approval across Table-to-Text Serialization Approaches

EMNLP 2025

Large Language Models (LLMs) are increasingly employed in high-stakes decision-making tasks, such as loan approvals. While their applications expand across domains, LLMs struggle to process tabular data, ensuring fairness and delivering reliable predictions. In this work, we assess the performance a

Cited by 0SourcePDFScholar
2025

Attention on Multiword Expressions: A Multilingual Study of BERT-based Models with Regard to Idiomaticity and Microsyntax

NAACL 2025findings

This study analyzes the attention patterns of fine-tuned encoder-only models based on the BERT architecture (BERT-based models) towards two distinct types of Multiword Expressions (MWEs): idioms and microsyntactic units (MSUs). Idioms present challenges in semantic non-compositionality, whereas MSUs…

2025

Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead

EMNLP 2025

With over 2,000 languages and potentially millions of speakers, Africa represents one of the richest linguistic regions in the world. Yet, this diversity is scarcely reflected in state-of-the-art natural language processing (NLP) systems and large language models (LLMs), which predominantly support

2025

Evaluating the Capabilities of Large Language Models for Multi-label Emotion Understanding

COLING 2025main

Large Language Models (LLMs) show promising learning and reasoning abilities. Compared to other NLP tasks, multilingual and multi-label emotion evaluation tasks are under-explored in LLMs. In this paper, we present EthioEmo, a multi-label emotion classification dataset for four Ethiopian languages,…

2025

INJONGO: A Multicultural Intent Detection and Slot-filling Dataset for 16 African Languages

ACL 2025long

Slot-filling and intent detection are well-established tasks in Conversational AI. However, current large-scale benchmarks for these tasks often exclude evaluations of low-resource languages and rely on translations from English benchmarks, thereby predominantly reflecting Western-centric concepts.…

Cited by 0SourcePDFScholar
2025

Improving Semantic Understanding in Speech Language Models via Brain-tuning

ICLR 2025poster

Speech language models align with human brain responses to natural language to an impressive degree. However, current models rely heavily on low-level speech features, indicating they lack brain-relevant semantics which limits their utility as model organisms of semantic processing in the brain. In…

2025

It’s Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems

ACL 2025long

Idioms are defined as a group of words with a figurative meaning not deducible from their individual components. Although modern machine translation systems have made remarkable progress, translating idioms remains a major challenge, especially for speech-to-text systems, where research on this topi…

2025

PricingLogic: Evaluating LLMs Reasoning on Complex Tourism Pricing Tasks

EMNLP 2025

We present PricingLogic, the first benchmarkthat probes whether Large Language Mod-els (LLMs) can reliably automate tourism-booking prices when multiple, overlapping farerules apply. Travel agencies are eager to of-fload this error-prone task to AI systems; how-ever, deploying LLMs without verified

2025

Unveiling the Key Factors for Distilling Chain-of-Thought Reasoning

ACL 2025finding

Large Language Models (LLMs) excel in reasoning tasks through Chain-of-Thought (CoT) prompting. However, CoT prompting greatly increases computational demands, which has prompted growing interest in distilling CoT capabilities into Small Language Models (SLMs). This study systematically examines the…

2024

A Preference-driven Paradigm for Enhanced Translation with Large Language Models

NAACL 2024long

Recent research has shown that large language models (LLMs) can achieve remarkable translation performance through supervised fine-tuning (SFT) using only a small amount of parallel data. However, SFT simply instructs the model to imitate the reference translations at the token level, making it vuln…

2024

Annotating Customer-Oriented Behaviour in Call Centre Sales Dialogues

COLING 2024main

Customer-oriented behaviour (COB) plays an important role in call centre interactions, particularly in the context of successful sales negotiation. However, the evaluation of COB in customer-agent conversations often lacks clarity in its definition and robust computational assessment methods. This p…

Cited by 1SourcePDFScholar
2024

EthioLLM: Multilingual Large Language Models for Ethiopian Languages with Task Evaluation

COLING 2024main

Large language models (LLMs) have gained popularity recently due to their outstanding performance in various downstream Natural Language Processing (NLP) tasks. However, low-resource languages are still lagging behind current state-of-the-art (SOTA) developments in the field of NLP due to insufficie…

2024

Fine-Tuning Large Language Models to Translate: Will a Touch of Noisy Data in Misaligned Languages Suffice?

EMNLP 2024main

Traditionally, success in multilingual machine translation can be attributed to three key factors in training data: large volume, diverse translation directions, and high quality. In the current practice of fine-tuning large language models (LLMs) for translation, we revisit the importance of these…

2024

From Insights to Actions: The Impact of Interpretability and Analysis Research on NLP

EMNLP 2024main

Interpretability and analysis (IA) research is a growing subfield within NLP with the goal of developing a deeper understanding of the behavior or inner workings of NLP systems and methods. Despite growing interest in the subfield, a criticism of this work is that it lacks actionable insights and th…

2024

Self-Supervised Adaptive Pre-Training of Multilingual Speech Models for Language and Dialect Identification

ICASSP 2024accepted

Transformer-based, pre-trained speech models have shown striking performance when fine-tuned on various downstream tasks such as automatic speech recognition and spoken language identification (SLID). However, the problem of domain mismatch remains a challenge in this area, where the domain of the p…

Cited by 0SourceScholar
2024

The Hidden Space of Transformer Language Adapters

ACL 2024long

We analyze the operation of transformer language adapters, which are small modules trained on top of a frozen language model to adapt its predictions to new target languages. We show that adapted predictions mostly evolve in the source language the model was trained on, while the target language bec…

2024

The Impact of Demonstrations on Multilingual In-Context Learning: A Multidimensional Analysis

ACL 2024findings

In-context learning is a popular inference strategy where large language models solve a task using only a few labeled demonstrations without needing any parameter updates. Although there have been extensive studies on English in-context learning, multilingual in-context learning remains under-explor…

2024

Understanding “Democratization” in NLP and ML Research

EMNLP 2024main

Recent improvements in natural language processing (NLP) and machine learning (ML) and increased mainstream adoption have led to researchers frequently discussing the “democratization” of artificial intelligence. In this paper, we seek to clarify how democratization is understood in NLP and ML publi…

2024

Who Did You Blame When Your Project Failed? Designing a Corpus for Presupposition Generation in Cross-Examination Dialogues

COLING 2024main

This paper introduces the corpus for the novel task of presupposition generation - a natural language generation problem where a model produces a list of presuppositions carried by the given input sentence, in the context of the presented research - given the cross-examination question. Two datasets…

Cited by 0SourcePDFScholar
2023

A Lightweight Method to Generate Unanswerable Questions in English

EMNLP 2023short findings

If a question cannot be answered with the available information, robust systems for question answering (QA) should know *not* to answer. One way to build QA models that do this is with additional training data comprised of unanswerable questions, created either by employing annotators or through aut…

Cited by 0SourcecodeScholar
2023

Few-shot Fine-tuning vs. In-context Learning: A Fair Comparison and Evaluation

ACL 2023findings

Few-shot fine-tuning and in-context learning are two alternative strategies for task adaptation of pre-trained language models. Recently, in-context learning has gained popularity over fine-tuning due to its simplicity and improved out-of-domain generalization, and because extensive evidence shows t…

2023

MasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African languages

ACL 2023long

In this paper, we present AfricaPOS, the largest part-of-speech (POS) dataset for 20 typologically diverse African languages. We discuss the challenges in annotating POS for these languages using the universal dependencies (UD) guidelines. We conducted extensive POS baseline experiments using both c…

2023

Weaker Than You Think: A Critical Look at Weakly Supervised Learning

ACL 2023long

Weakly supervised learning is a popular approach for training machine learning models in low-resource settings. Instead of requesting high-quality yet costly human annotations, it allows training models with noisy annotations obtained from various weak sources. Recently, many sophisticated approache…

2022

A Few Thousand Translations Go a Long Way! Leveraging Pre-trained Models for African News Translation

NAACL 2022long

Recent advances in the pre-training for language models leverage large-scale datasets to create multilingual models. However, low-resource languages are mostly left out in these datasets. This is primarily because many widely spoken languages that are not well represented on the web and therefore ex…

2022

Adapting Pre-trained Language Models to African Languages via Multilingual Adaptive Fine-Tuning

COLING 2022main

Multilingual pre-trained language models (PLMs) have demonstrated impressive performance on several downstream tasks for both high-resourced and low-resourced languages. However, there is still a large performance drop for languages unseen during pre-training, especially African languages. One of th…

2022

Call-Sign Recognition and Understanding for Noisy Air-Traffic Transcripts Using Surveillance Information

ICASSP 2022accepted

Air traffic control (ATC) relies on communication via speech between pilot and air-traffic controller (ATCO). The call-sign, as unique identifier for each flight, is used to address a specific pilot by the ATCO. Extracting the call-sign from the communication is a challenge because of the noisy ATC…

Cited by 0SourceScholar
2022

Label-Descriptive Patterns and Their Application to Characterizing Classification Errors

ICML 2022spotlight

State-of-the-art deep learning methods achieve human-like performance on many tasks, but make errors nevertheless. Characterizing these errors in easily interpretable terms gives insight into whether a classifier is prone to making systematic errors, but also gives a way to act and improve the class…

2022

MCSE: Multimodal Contrastive Learning of Sentence Embeddings

NAACL 2022long

Learning semantically meaningful sentence embeddings is an open problem in natural language processing. In this work, we propose a sentence embedding learning approach that exploits both visual and textual information via a multimodal contrastive objective. Through experiments on a variety of semant…

2022

MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition

EMNLP 2022main

African languages are spoken by over a billion people, but they are under-represented in NLP research and development. Multiple challenges exist, including the limited availability of annotated training and evaluation datasets as well as the lack of understanding of which settings, languages, and re…

2021

A Survey on Recent Approaches for Natural Language Processing in Low-Resource Scenarios

NAACL 2021long

Deep neural networks and huge language models are becoming omnipresent in natural language applications. As they are known for requiring large amounts of training data, there is a growing body of work to improve the performance in low-resource settings. Motivated by the recent fundamental changes to…

Cited by 393SourcePDFScholar
2021

Analysing the Noise Model Error for Realistic Noisy Label Data

AAAI 2021technical

Distant and weak supervision allow to obtain large amounts of labeled training data quickly and cheaply, but these automatic annotations tend to contain a high amount of errors. A popular technique to overcome the negative effects of these noisy labels is noise modelling where the underlying noise p…

2021

FAME: Feature-Based Adversarial Meta-Embeddings for Robust Input Representations

EMNLP 2021main

Combining several embeddings typically improves performance in downstream tasks as different embeddings encode different information. It has been shown that even models using embeddings from transformers still benefit from the inclusion of standard word embeddings. However, the combination of embedd…

2021

On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong Baselines

ICLR 2021poster

Fine-tuning pre-trained transformer-based language models such as BERT has become a common practice dominating leaderboards across various NLP benchmarks. Despite the strong empirical performance of fine-tuned models, fine-tuning is an unstable process: training the same model with multiple random s…

2021

Preventing Author Profiling through Zero-Shot Multilingual Back-Translation

EMNLP 2021main

Documents as short as a single sentence may inadvertently reveal sensitive information about their authors, including e.g. their gender or ethnicity. Style transfer is an effective way of transforming texts in order to remove any information that enables author profiling. However, for a number of cu…

2021

To Share or not to Share: Predicting Sets of Sources for Model Transfer Learning

EMNLP 2021main

In low-resource settings, model transfer can help to overcome a lack of labeled data for many tasks and domains. However, predicting useful transfer sources is a challenging problem, as even the most similar sources might lead to unexpected negative transfer results. Thus, ranking methods based on t…

2020

A Closer Look at Linguistic Knowledge in Masked Language Models: The Case of Relative Clauses in American English

COLING 2020main

Transformer-based language models achieve high performance on various tasks, but we still lack understanding of the kind of linguistic knowledge they learn and rely on. We evaluate three models (BERT, RoBERTa, and ALBERT), testing their grammatical and semantic knowledge by sentence-level probing, d…