← Search

Lirong Qiu

4 accepted papers

2026

MirrorShield: Towards Dynamic Adaptive Defense Against Jailbreaks via Entropy-Guided Mirror Crafting

AAAI 2026technical

Defending large language models (LLMs) against jailbreak attacks is crucial for ensuring their safe deployment. Existing defense strategies typically rely on predefined static criteria to differentiate between harmful and benign prompts. However, such rigid rules fail to accommodate the inherent com

Cited by 0SourcePDFScholar
2025

Feint and Attack: Jailbreaking and Protecting LLMs via Attention Distribution Modeling

IJCAI 2025

Most jailbreak methods for large language models (LLMs) focus on superficially improving attack success through manually defined rules. However, they fail to uncover the underlying mechanisms within target LLMs that explain why an attack succeeds or fails. In this paper, we propose investigating the

Cited by 0SourcePDFScholar
2025

OSTAR: Optimized Statistical Text-classifier with Adversarial Resistance

NeurIPS 2025poster

The advancements in generative models and the real-world attack of machine-generated text(MGT) create a demand for more robust detection methods. The existing MGT detection methods for adversarial environments primarily consist of manually designed statistical-based methods and fine-tuned classifi…

Cited by 0SourcecodeScholar
2024

BaitAttack: Alleviating Intention Shift in Jailbreak Attacks via Adaptive Bait Crafting

EMNLP 2024main

Jailbreak attacks enable malicious queries to evade detection by LLMs. Existing attacks focus on meticulously constructing prompts to disguise harmful intentions. However, the incorporation of sophisticated disguising prompts may incur the challenge of “intention shift”. Intention shift occurs when…

Cited by 3SourcePDFScholar