2025
MixAT: Combining Continuous and Discrete Adversarial Training for LLMs
NeurIPS 2025poster
Despite recent efforts in Large Language Model (LLM) safety and alignment, current adversarial attacks on frontier LLMs can still consistently force harmful generations. Although adversarial training has been widely studied and shown to significantly improve the robustness of traditional machine le…