← Search

Baolei Zhang

2 accepted papers

2025

Prompt-Guided Internal States for Hallucination Detection of Large Language Models

ACL 2025long

Large Language Models (LLMs) have demonstrated remarkable capabilities across a variety of tasks in different domains. However, they sometimes generate responses that are logically coherent but factually incorrect or misleading, which is known as LLM hallucinations. Data-driven supervised methods tr…

2024

BadActs: A Universal Backdoor Defense in the Activation Space

ACL 2024findings

Backdoor attacks pose an increasingly severe security threat to Deep Neural Networks (DNNs) during their development stage. In response, backdoor sample purification has emerged as a promising defense mechanism, aiming to eliminate backdoor triggers while preserving the integrity of the clean conten…