← Search

Depeng Xu

7 accepted papers

2026

Broadening the Backdoor Basin: Understanding LLM Backdoors Collapse and Making Backdoors Persistent

ICML 2026poster

Large Language Models (LLMs) have shown to be vulnerable to backdoor attacks, yet we observe that many LLM backdoors do not survive when end users perform supervised fine-tuning (SFT). In this work, we provide a geometric explanation: by probing the backdoor objective under controlled weight perturb…

Cited by 0SourceScholar
2026

Don't Shift the Trigger: Robust Gradient Ascent for Backdoor Unlearning

ICLR 2026poster

Backdoor attacks pose a significant threat to machine learning models, allowing adversaries to implant hidden triggers that alter model behavior when activated. Although gradient ascent (GA)-based unlearning has been proposed as an efficient backdoor removal approach, we identify a critical yet over…

Cited by 0SourceScholar
2025

Fine-tuning LLMs with Cross-Attention-based Weight Decay for Bias Mitigation

EMNLP 2025

Large Language Models (LLMs) excel in Natural Language Processing (NLP) tasks but often propagate societal biases from their training data, leading to discriminatory outputs. These biases are amplified by the models’ self-attention mechanisms, which disproportionately emphasize biased correlations w

2025

Privacy-centric Deep Motion Retargeting for Anonymization of Skeleton-Based Motion Visualization

ICCV 2025poster

Capturing and visualizing motion using skeleton-based techniques is a key aspect of computer vision, particularly in virtual reality (VR) settings. Its popularity has surged, driven by the simplicity of obtaining skeleton data and the growing appetite for virtual interaction. Although this skeleton…

2024

Defense against Backdoor Attack on Pre-trained Language Models via Head Pruning and Attention Normalization

ICML 2024poster

Pre-trained language models (PLMs) are commonly used for various downstream natural language processing tasks via fine-tuning. However, recent studies have demonstrated that PLMs are vulnerable to backdoor attacks, which can mislabel poisoned samples to target outputs even after a vanilla fine-tunin…

Cited by 5SourcePDFScholar
2024

Discovering and Mitigating Indirect Bias in Attention-Based Model Explanations

NAACL 2024findings

As the field of Natural Language Processing (NLP) increasingly adopts transformer-based models, the issue of bias becomes more pronounced. Such bias, manifesting through stereotypes and discriminatory practices, can disadvantage certain groups. Our study focuses on direct and indirect bias in the mo…