SRD: Reinforcement-Learned Semantic Perturbation for Backdoor Defense in VLMs
Visual language models (VLMs) have made significant progress in image captioning tasks, yet recent studies have found they are vulnerable to backdoor attacks. Attackers can inject undetectable perturbations into the data during inference, triggering abnormal behavior and generating malicious caption