AAAI 2026technical0 citations

AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs

Debdeep Sanyal, Manodeep Ray, Murari Mandal

Abstract

The release of open-weight large language models (LLMs) creates a tension between advancing accessible research and preventing misuse, such as malicious fine-tuning to elicit harmful content. Current safety measures struggle to preserve the general capabilities of the LLM while resisting a determined adversary with full access to the model

BibTeX
@inproceedings{aaai2026_antidotebilevela,
  title = {AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs},
  author = {Debdeep Sanyal and Manodeep Ray and Murari Mandal},
  booktitle = {AAAI 2026},
  year = {2026}
}
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs · AAAI 2026