← Search

Heqing Huang

2 accepted papers

2026

Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks

ICML 2026poster

As Large Language Model (LLM) agents are increasingly deployed in complex environments, multi-turn interaction attacks have become a significant security challenge. Existing detection methods typically rely on historical context. However, this retrospective logic struggles to identify deep malicious…

Cited by 0SourceScholar
2025

The Illusion of Role Separation: Hidden Shortcuts in LLM Role Learning (and How to Fix Them)

ICML 2025poster

Large language models (LLMs) that integrate multiple input roles (e.g., system instructions, user queries, external tool outputs) are increasingly prevalent in practice. Ensuring that the model accurately distinguishes messages from each role—a concept we call *role separation*—is crucial for consis…

Cited by 0SourcePDFScholar