← Search

Greta Warren

2 accepted papers

2026

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

ICML 2026poster

Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work has shown that this increase in capability comes with a cost: it can increase a model's tendency to respond to unsafe adversarial prompts, even when fine-tuning w…

Cited by 0SourceScholar
2025

Can Community Notes Replace Professional Fact-Checkers?

ACL 2025short

Two commonly employed strategies to combat the rise of misinformation on social media are (i) fact-checking by professional organisations and (ii) community moderation by platform users. Policy changes by Twitter/X and, more recently, Meta, signal a shift away from partnerships with fact-checking or…