← Search

Shuhan Yuan

10 accepted papers

2026

Broadening the Backdoor Basin: Understanding LLM Backdoors Collapse and Making Backdoors Persistent

ICML 2026poster

Large Language Models (LLMs) have shown to be vulnerable to backdoor attacks, yet we observe that many LLM backdoors do not survive when end users perform supervised fine-tuning (SFT). In this work, we provide a geometric explanation: by probing the backdoor objective under controlled weight perturb…

Cited by 0SourceScholar
2026

Don't Shift the Trigger: Robust Gradient Ascent for Backdoor Unlearning

ICLR 2026poster

Backdoor attacks pose a significant threat to machine learning models, allowing adversaries to implant hidden triggers that alter model behavior when activated. Although gradient ascent (GA)-based unlearning has been proposed as an efficient backdoor removal approach, we identify a critical yet over…

Cited by 0SourceScholar
2025

Fine-tuning LLMs with Cross-Attention-based Weight Decay for Bias Mitigation

EMNLP 2025

Large Language Models (LLMs) excel in Natural Language Processing (NLP) tasks but often propagate societal biases from their training data, leading to discriminatory outputs. These biases are amplified by the models’ self-attention mechanisms, which disproportionately emphasize biased correlations w

2025

Privacy-centric Deep Motion Retargeting for Anonymization of Skeleton-Based Motion Visualization

ICCV 2025poster

Capturing and visualizing motion using skeleton-based techniques is a key aspect of computer vision, particularly in virtual reality (VR) settings. Its popularity has surged, driven by the simplicity of obtaining skeleton data and the growing appetite for virtual interaction. Although this skeleton…

2025

Root Cause Analysis of Anomalies in Multivariate Time Series through Granger Causal Discovery

ICLR 2025oral

Identifying the root causes of anomalies in multivariate time series is challenging due to the complex dependencies among the series. In this paper, we propose a comprehensive approach called AERCA that inherently integrates Granger causal discovery with root cause analysis. By defining anomalies as…

Cited by 0SourcePDFScholar
2024

Defense against Backdoor Attack on Pre-trained Language Models via Head Pruning and Attention Normalization

ICML 2024poster

Pre-trained language models (PLMs) are commonly used for various downstream natural language processing tasks via fine-tuning. However, recent studies have demonstrated that PLMs are vulnerable to backdoor attacks, which can mislabel poisoned samples to target outputs even after a vanilla fine-tunin…

Cited by 5SourcePDFScholar
2024

Discovering and Mitigating Indirect Bias in Attention-Based Model Explanations

NAACL 2024findings

As the field of Natural Language Processing (NLP) increasingly adopts transformer-based models, the issue of bias becomes more pronounced. Such bias, manifesting through stereotypes and discriminatory practices, can disadvantage certain groups. Our study focuses on direct and indirect bias in the mo…

2022

On the Convergence of the Monte Carlo Exploring Starts Algorithm for Reinforcement Learning

ICLR 2022poster

A simple and natural algorithm for reinforcement learning (RL) is Monte Carlo Exploring Starts (MCES), where the Q-function is estimated by averaging the Monte Carlo returns, and the policy is improved by choosing actions that maximize the current estimate of the Q-function. Exploration is performed…

Cited by 28SourcePDFScholar
2022

Robust Unstructured Knowledge Access in Conversational Dialogue with ASR Errors

ICASSP 2022accepted

Performance of spoken language understanding (SLU) can be degraded with automatic speech recognition (ASR) errors. We propose a novel approach to improve SLU robustness by randomly corrupting clean training text with an ASR error simulator, followed by self-correcting the errors and minimizing the t…

Cited by 0SourceScholar