← Search

Zitao Xuan

1 accepted papers

2025

ShieldHead: Decoding-time Safeguard for Large Language Models

ACL 2025finding

In light of the widespread deployment of Large Language Models (LLMs), the responsibility for safeguarding and regulating LLM-generated content has taken on heightened significance. Recent advancements in LLM-based moderation methods, e.g., LlamaGuard, have demonstrated remarkable promise in identif…