AAAI 2026technical0 citations
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
Debdeep Sanyal, Manodeep Ray, Murari Mandal
Abstract
The release of open-weight large language models (LLMs) creates a tension between advancing accessible research and preventing misuse, such as malicious fine-tuning to elicit harmful content. Current safety measures struggle to preserve the general capabilities of the LLM while resisting a determined adversary with full access to the model
BibTeX
@inproceedings{aaai2026_antidotebilevela,
title = {AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs},
author = {Debdeep Sanyal and Manodeep Ray and Murari Mandal},
booktitle = {AAAI 2026},
year = {2026}
}