← Search

Lea Frermann

16 accepted papers

2026

Control Illusion: The Failure of Instruction Hierarchies in Large Language Models

AAAI 2026technical

Large language models (LLMs) are increasingly deployed with hierarchical instruction schemes, where certain instructions (e.g., system-level directives) are expected to take precedence over others (e.g., user messages). Yet, we lack a systematic understanding of how effectively these hierarchical co

Cited by 0SourcePDFScholar
2026

Multilingual Unlearning in LLMs: Transfer, Dynamics, and Reversibility

ICML 2026poster

Large language models (LLMs) can memorize sensitive facts, motivating *unlearning* methods that remove targeted knowledge without costly retraining. However, unlearning research remains heavily English-centric. We study multilingual unlearning by extending the TOFU benchmark to five languages, and f…

Cited by 0SourceScholar
2025

Comparing Moral Values in Western English-speaking societies and LLMs with Word Associations

ACL 2025long

As the impact of large language models increases, understanding the moral values they encode becomes ever more important. Assessing moral values encoded in these models via direct prompting is challenging due to potential leakage of human norms into model training data, and their sensitivity to prom…

2025

Human Interest Framing across Cultures: A Case Study on Climate Change

COLING 2025main

Human Interest (HI) framing is a narrative strategy that injects news stories with a relatable, emotional angle and a human face to engage the audience. In this study we investigate the use of HI framing across different English-speaking cultures in news articles about climate change. Despite its de…

Cited by 0SourcePDFScholar
2025

Moderation Matters: Measuring Conversational Moderation Impact in English as a Second Language Group Discussion

ACL 2025finding

English as a Second Language (ESL) speakers often struggle to engage in group discussions due to language barriers. While moderators can facilitate participation, few studies assess conversational engagement and evaluate moderation effectiveness. To address this gap, we develop a dataset comprising…

2025

WHoW: A Cross-domain Approach for Analysing Conversation Moderation

NAACL 2025long

We propose WHoW, an evaluation framework for analyzing the facilitation strategies of moderators across different domains/scenarios by examining their motives (Why), dialogue acts (How) and target speaker (Who). Using this framework, we annotated 5,657 moderation sentences with human judges and 15,4…

2024

Media Framing: A typology and Survey of Computational Approaches Across Disciplines

ACL 2024long

Framing studies how individuals and societies make sense of the world, by communicating or representing complex issues through schema of interpretation. The framing of information in the mass media influences our interpretation of facts and corresponding decisions, so detecting and analysing it is e…

2023

Conflicts, Villains, Resolutions: Towards models of Narrative Media Framing

ACL 2023long

Despite increasing interest in the automatic detection of media frames in NLP, the problem is typically simplified as single-label classification and adopts a topic-like view on frames, evading modelling the broader document-level narrative. In this work, we revisit a widely used conceptualization o…

2023

More than Votes? Voting and Language based Partisanship in the US Supreme Court

EMNLP 2023short findings

Understanding the prevalence and dynamics of justice partisanship and ideology in the US Supreme Court is critical in studying jurisdiction. Most research quantifies partisanship based on voting behavior, and oral arguments in the courtroom --- the last essential procedure before the final case outc…

Cited by 0SourceScholar
2022

A Computational Acquisition Model for Multimodal Word Categorization

NAACL 2022long

Recent advances in self-supervised modeling of text and images open new opportunities for computational models of child language acquisition, which is believed to rely heavily on cross-modal signals. However, prior studies has been limited by their reliance on vision models trained on large image da…

2022

Optimising Equal Opportunity Fairness in Model Training

NAACL 2022long

Real-world datasets often encode stereotypes and societal biases. Such biases can be implicitly captured by trained models, leading to biased predictions and exacerbating existing societal preconceptions. Existing debiasing methods, such as adversarial training and removing protected information fro…

2022

Unsupervised Cross-Lingual Transfer of Structured Predictors without Source Data

NAACL 2022long

Providing technologies to communities or domains where training data is scarce or protected e.g., for privacy reasons, is becoming increasingly important. To that end, we generalise methods for unsupervised transfer from multiple input models for structured prediction. We show that the means of aggr…

2021

Evaluating Debiasing Techniques for Intersectional Biases

EMNLP 2021main

Bias is pervasive for NLP models, motivating the development of automatic debiasing techniques. Evaluation of NLP debiasing methods has largely been limited to binary attributes in isolation, e.g., debiasing with respect to binary gender or race, however many corpora involve multiple such attributes…

Cited by 55SourcePDFScholar
2021

Fairness-aware Class Imbalanced Learning

EMNLP 2021main

Class imbalance is a common challenge in many NLP tasks, and has clear connections to bias, in that bias in training data often leads to higher accuracy for majority groups at the expense of minority groups. However there has traditionally been a disconnect between research on class-imbalanced learn…

Cited by 35SourcePDFScholar
2021

Framing Unpacked: A Semi-Supervised Interpretable Multi-View Model of Media Frames

NAACL 2021long

Understanding how news media frame political issues is important due to its impact on public attitudes, yet hard to automate. Computational approaches have largely focused on classifying the frame of a full news article while framing signals are often subtle and local. Furthermore, automatic news an…