← Search

Katja Filippova

6 accepted papers

2025

Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research

NeurIPS 2025oral

"Machine unlearning" is a popular proposed solution for mitigating the existence of content in an AI model that is problematic for legal or moral reasons, including privacy, copyright, safety, and more. For example, unlearning is often invoked as a solution for removing the effects of specific infor…

Cited by 0SourceScholar
2023

Dissecting Recall of Factual Associations in Auto-Regressive Language Models

EMNLP 2023long main

Transformer-based language models (LMs) are known to capture factual knowledge in their parameters. While previous work looked into where factual associations are stored, only little is known about how they are retrieved internally during inference. We investigate this question through the lens of i…

Cited by 0SourceScholar
2023

Make Every Example Count: On the Stability and Utility of Self-Influence for Learning from Noisy NLP Datasets

EMNLP 2023long main

Increasingly larger datasets have become a standard ingredient to advancing the state-of-the-art in NLP. However, data quality might have already become the bottleneck to unlock further gains. Given the diversity and the sizes of modern datasets, standard data filtering is not straight-forward to ap…

Cited by 0SourceScholar
2023

Theoretical and Practical Perspectives on what Influence Functions Do

NeurIPS 2023spotlight

Influence functions (IF) have been seen as a technique for explaining model predictions through the lens of the training data. Their utility is assumed to be in identifying training examples "responsible" for a prediction so that, for example, correcting a prediction is possible by intervening on th…

Cited by 22SourcePDFScholar
2022

“Will You Find These Shortcuts?” A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification

EMNLP 2022main

Feature attribution a.k.a. input salience methods which assign an importance score to a feature are abundant but may produce surprisingly different results for the same model on the same input. While differences are expected if disparate definitions of importance are assumed, most methods claim to p…

Cited by 69SourcePDFScholar
2021

Controlling Machine Translation for Multiple Attributes with Additive Interventions

EMNLP 2021main

Fine-grained control of machine translation (MT) outputs along multiple attributes is critical for many modern MT applications and is a requirement for gaining users’ trust. A standard approach for exerting control in MT is to prepend the input with a special tag to signal the desired output attribu…

Cited by 30SourcePDFScholar