2025
ICLScan: Detecting Backdoors in Black-Box Large Language Models via Targeted In-context Illumination
NeurIPS 2025poster
The widespread deployment of large language models (LLMs) allows users to access their capabilities via black-box APIs, but backdoor attacks pose serious security risks for API users by hijacking the model behavior. This highlights the importance of backdoor detection technologies to help users audi…