2026
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
AAAI 2026technical
The release of open-weight large language models (LLMs) creates a tension between advancing accessible research and preventing misuse, such as malicious fine-tuning to elicit harmful content. Current safety measures struggle to preserve the general capabilities of the LLM while resisting a determine