← Search

Paul Youssef

6 accepted papers

2025

Has this Fact been Edited? Detecting Knowledge Edits in Language Models

NAACL 2025long

Knowledge editing methods (KEs) can update language models’ obsolete or inaccurate knowledge learned from pre-training. However, KEs can be used for malicious applications, e.g., inserting misinformation and toxic content. Knowing whether a generated output is based on edited knowledge or first-hand…

2025

How to Make LLMs Forget: On Reversing In-Context Knowledge Edits

NAACL 2025long

In-context knowledge editing (IKE) enables efficient modification of large language model (LLM) outputs without parameter changes and at zero-cost. However, it can be misused to manipulate responses opaquely, e.g., insert misinformation or offensive content. Such malicious interventions could be inc…

2025

Position: Editing Large Language Models Poses Serious Safety Risks

ICML 2025poster

Large Language Models (LLMs) contain large amounts of facts about the world. These facts can become outdated over time, which has led to the development of knowledge editing methods (KEs) that can change specific facts in LLMs with limited side effects. This position paper argues that editing LLMs p…

Cited by 3SourcePDFScholar
2024

LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study

EMNLP 2024finding

As NLP models become more complex, understanding their decisions becomes more crucial. Counterfactuals (CFs), where minimal changes to inputs flip a model’s prediction, offer a way to explain these models. While Large Language Models (LLMs) have shown remarkable performance in NLP tasks, their effic…

2023

Give Me the Facts! A Survey on Factual Knowledge Probing in Pre-trained Language Models

EMNLP 2023long findings

Pre-trained Language Models (PLMs) are trained on vast unlabeled data, rich in world knowledge. This fact has sparked the interest of the community in quantifying the amount of factual knowledge present in PLMs, as this explains their performance on downstream tasks, and potentially justifies their…

Cited by 37SourceScholar