2025
LLMScan: Causal Scan for LLM Misbehavior Detection
ICML 2025poster
Despite the success of Large Language Models (LLMs) across various fields, their potential to generate untruthful and harmful responses poses significant risks, particularly in critical applications. This highlights the urgent need for systematic methods to detect and prevent such misbehavior. While…