← Search

Seraphina Goldfarb-Tarrant

10 accepted papers

2026

Learning is Forgetting; LLM Training As Lossy Compression

ICLR 2026poster

Despite the increasing prevalence of large language models (LLMs), we still have a limited understanding of how their representational spaces are structured. This limits our ability to interpret how and what they learn or relate them to learning in humans. We argue LLMs are best seen as an instance…

Cited by 0SourcecodeScholar
2025

A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users

EMNLP 2025

To assist users in complex tasks, LLMs generate plans: step-by-step instructions towards a goal. While alignment methods aim to ensure LLM plans are helpful, they train (RLHF) or evaluate (ChatbotArena) on what users prefer, assuming this reflects what helps them. We test this with Planorama: an int

Cited by 0SourcePDFScholar
2024

A SMART Mnemonic Sounds like “Glue Tonic”: Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick

EMNLP 2024main

Keyword mnemonics are memorable explanations that link new terms to simpler keywords.Prior work generates mnemonics for students, but they do not train models using mnemonics students prefer and aid learning.We build SMART, a mnemonic generator trained on feedback from real students learning new ter…

2024

The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm

EMNLP 2024main

A key concern with the concept of *“alignment”* is the implicit question of *“alignment to what?”*. AI systems are increasingly used across the world, yet safety alignment is often focused on homogeneous monolingual settings. Additionally, preference training and safety measures often overfit to har…

Cited by 18SourcePDFScholar
2023

Bias Beyond English: Counterfactual Tests for Bias in Sentiment Analysis in Four Languages

ACL 2023findings

Sentiment analysis (SA) systems are used in many products and hundreds of languages. Gender and racial biases are well-studied in English SA systems, but understudied in other languages, with few resources for such studies. To remedy this, we build a counterfactual evaluation corpus for gender and r…

Cited by 17SourcePDFScholar
2023

This prompt is measuring <mask>: evaluating bias evaluation in language models

ACL 2023findings

Bias research in NLP seeks to analyse models for social biases, thus helping NLP practitioners uncover, measure, and mitigate social harms. We analyse the body of work that uses prompts and templates to assess bias in language models. We draw on a measurement modelling framework to create a taxonomy…

Cited by 34SourcePDFScholar
2022

How Gender Debiasing Affects Internal Model Representations, and Why It Matters

NAACL 2022long

Common studies of gender bias in NLP focus either on extrinsic bias measured by model performance on a downstream task or on intrinsic bias found in models’ internal representations. However, the relationship between extrinsic and intrinsic bias is relatively unknown. In this work, we illuminate thi…

Cited by 37SourcePDFScholar
2021

Intrinsic Bias Metrics Do Not Correlate with Application Bias

ACL 2021long

Natural Language Processing (NLP) systems learn harmful societal biases that cause them to amplify inequality as they are deployed in more and more situations. To guide efforts at debiasing these systems, the NLP community relies on a variety of metrics that quantify bias in models. Some of these me…

Cited by 188SourcePDFScholar