← Search

Leon Weber-Genzel

3 accepted papers

2024

VariErr NLI: Separating Annotation Error from Human Label Variation

ACL 2024long

Human label variation arises when annotators assign different labels to the same item for valid reasons, while annotation errors occur when labels are assigned for invalid reasons. These two issues are prevalent in NLP benchmarks, yet existing research has studied them in isolation. To the best of o…

2024

“My Answer is C”: First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models

ACL 2024findings

The open-ended nature of language generation makes the evaluation of autoregressive large language models (LLMs) challenging. One common evaluation approach uses multiple-choice questions to limit the response space. The model is then evaluated by ranking the candidate answers by the log probability…

2023

Establishing Trustworthiness: Rethinking Tasks and Model Evaluation

EMNLP 2023short main

Language understanding is a multi-faceted cognitive capability, which the Natural Language Processing (NLP) community has striven to model computationally for decades. Traditionally, facets of linguistic intelligence have been compartmentalized into tasks with specialized model architectures and cor…

Cited by 0SourceScholar