AdvBDGen: A Robust Framework for Generating Adaptive and Stealthy Backdoors in LLM Alignment
With the increasing adoption of reinforcement learning with human feedback (RLHF) to align large language models (LLMs), the risk of backdoor installation during the alignment process has grown, potentially leading to unintended and harmful behaviors. Existing backdoor attacks mostly focus on simple