AAAI 2026technical0 citations
ShadeEdit: A Utility-Preserving and Defense-Evasive Knowledge Manipulation Attack in Federated LLMs
Xu Zhang, Hangcheng Liu, Shangwei Guo, Shudong Zhang, Tianwei Zhang, Tao Xiang
Abstract
Recent studies reveal that adversaries can manipulate the internal knowledge of large language models (LLMs) on selected topics through model editing, causing attacker-specified harmful or biased outputs when queried about the edited content. Once such tampered LLMs are distributed, they can mislead users on the targeted topics, thereby potentially propagating misinformation or reinforcing stereotypes. However, existing knowledge manipulation attacks rely on the ability to redistribute compromised models, which is infeasible in constrained settings like Federated Instruction Tuning (FedIT), where a central server controls LLM
BibTeX
@inproceedings{aaai2026_shadeeditautilit,
title = {ShadeEdit: A Utility-Preserving and Defense-Evasive Knowledge Manipulation Attack in Federated LLMs},
author = {Xu Zhang and Hangcheng Liu and Shangwei Guo and Shudong Zhang and Tianwei Zhang and Tao Xiang},
booktitle = {AAAI 2026},
year = {2026}
}