← Search

Xuehai Tang

8 accepted papers

2026

Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs

AAAI 2026technical

Large Language Models (LLMs) demonstrate impressive capabilities across diverse tasks, yet their safety mechanisms remain susceptible to adversarial exploitation of cognitive biases---systematic deviations from rational judgment. Unlike prior studies focusing on isolated biases, this work highlights

Cited by 0SourcePDFScholar
2026

Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment Through Latent Acoustic Pattern Triggers

AAAI 2026technical

As Audio Large Language Models (ALLMs) emerge as powerful tools for speech processing, their safety implications demand urgent attention. While considerable research has explored textual and vision safety, audio’s distinct characteristics present significant challenges. This paper first investigates

Cited by 0SourcePDFScholar
2026

Profiling the Irrational Agent: Cognitive Modeling of LLM Behaviors in Sequential Jailbreaks

ICML 2026poster

Large language models (LLMs) are increasingly deployed in high-stakes settings, yet they remain vulnerable to sequential jailbreaks that exploit multi-turn interaction to circumvent safety mechanisms. Current safety evaluations are largely outcome-based, offering little insight into the latent decis…

Cited by 0SourceScholar
2025

AdaPPA: Adaptive Position Pre-Fill Jailbreak Attack Approach Targeting LLMs

ICASSP 2025accepted

Jailbreak vulnerabilities in Large Language Models (LLMs) refer to methods that extract malicious content from the model by carefully crafting prompts or suffixes, which has garnered significant attention from the research community. However, traditional attack methods, which primarily focus on the…

Cited by 0SourceScholar
2025

Chain of Attack: Hide Your Intention through Multi-Turn Interrogation

ACL 2025finding

The latent knowledge of large language models (LLMs) contains harmful or unethical content, which introduces significant security risks upon their widespread deployment. Conducting jailbreak attacks on LLMs can proactively identify vulnerabilities to enhance their security measures. However, previou…

2025

Gamma-Guard: Lightweight Residual Adapters for Robust Guardrails in Large Language Models

EMNLP 2025

Large language models (LLMs) are widely deployed as zero-shot evaluators for answer grading, content moderation, and document ranking. Yet studies show that guard models (Guards)—LLMs fine-tuned for safety—remain vulnerable to “jailbreak” attacks, jeopardising downstream chatbots.We confirm this wea

2025

LyapLock: Bounded Knowledge Preservation in Sequential Large Language Model Editing

EMNLP 2025

Large Language Models often contain factually incorrect or outdated knowledge, giving rise to model editing methods for precise knowledge updates. However, current mainstream locate-then-edit approaches exhibit a progressive performance decline during sequential editing, due to inadequate mechanisms

2025

Segment-Recurrent Transformer with Multi-Scale Fusion for Long-Term Time Series Forecasting

ICASSP 2025accepted

Long-term time series forecasting (LTSF) seeks to make accurate long-term predictions by leveraging extensive historical data, which is crucial for solving scientific and engineering challenges. Traditional transformer-based methods process historical segments individually, leading to a limited view…

Cited by 0SourceScholar