← Search

Michael A. Hedderich

11 accepted papers

2026

Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering

ICLR 2026poster

Large vision-language models (VLMs) achieve strong performance in Visual Question Answering but still rely heavily on supervised fine-tuning (SFT) with massive labeled datasets, which is costly due to human annotations. Crucially, real-world datasets often exhibit *human uncertainty* (**HU**) — var…

Cited by 0SourceScholar
2025

Charting the Landscape of African NLP: Mapping Progress and Shaping the Road Ahead

EMNLP 2025

With over 2,000 languages and potentially millions of speakers, Africa represents one of the richest linguistic regions in the world. Yet, this diversity is scarcely reflected in state-of-the-art natural language processing (NLP) systems and large language models (LLMs), which predominantly support

2025

MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs

EMNLP 2025

Large language models (LLMs) are used globally across many languages, but their English-centric pretraining raises concerns about cross-lingual disparities for cultural awareness, often resulting in biased outputs. However, comprehensive multilingual evaluation remains challenging due to limited ben

2025

Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set

ACL 2025long

Discourse understanding is essential for many NLP tasks, yet most existing work remains constrained by framework-dependent discourse representations. This work investigates whether large language models (LLMs) capture discourse knowledge that generalizes across languages and frameworks. We address t…

2025

Semantic Component Analysis: Introducing Multi-Topic Distributions to Clustering-Based Topic Modeling

EMNLP 2025

Topic modeling is a key method in text analysis, but existing approaches fail to efficiently scale to large datasets or are limited by assuming one topic per document. Overcoming these limitations, we introduce Semantic Component Analysis (SCA), a topic modeling technique that discovers multiple top

2025

What’s the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns

ACL 2025long

Prompt engineering for large language models is challenging, as even small prompt perturbations or model changes can significantly impact the generated output texts. Existing evaluation methods of LLM outputs, either automated metrics or human evaluation, have limitations, such as providing limited…

2024

The Potential and Challenges of Evaluating Attitudes, Opinions, and Values in Large Language Models

EMNLP 2024finding

Recent advances in Large Language Models (LLMs) have sparked wide interest in validating and comprehending the human-like cognitive-behavioral traits LLMs may capture and convey. These cognitive-behavioral traits include typically Attitudes, Opinions, Values (AOVs). However, measuring AOVs embedded…

2022

Label-Descriptive Patterns and Their Application to Characterizing Classification Errors

ICML 2022spotlight

State-of-the-art deep learning methods achieve human-like performance on many tasks, but make errors nevertheless. Characterizing these errors in easily interpretable terms gives insight into whether a classifier is prone to making systematic errors, but also gives a way to act and improve the class…

2022

MCSE: Multimodal Contrastive Learning of Sentence Embeddings

NAACL 2022long

Learning semantically meaningful sentence embeddings is an open problem in natural language processing. In this work, we propose a sentence embedding learning approach that exploits both visual and textual information via a multimodal contrastive objective. Through experiments on a variety of semant…

2021

A Survey on Recent Approaches for Natural Language Processing in Low-Resource Scenarios

NAACL 2021long

Deep neural networks and huge language models are becoming omnipresent in natural language applications. As they are known for requiring large amounts of training data, there is a growing body of work to improve the performance in low-resource settings. Motivated by the recent fundamental changes to…

Cited by 393SourcePDFScholar
2021

Analysing the Noise Model Error for Realistic Noisy Label Data

AAAI 2021technical

Distant and weak supervision allow to obtain large amounts of labeled training data quickly and cheaply, but these automatic annotations tend to contain a high amount of errors. A popular technique to overcome the negative effects of these noisy labels is noise modelling where the underlying noise p…