Prompt Crossing: Evaluating Whether LLM Response Stem from Jailbreak or Normal Prompt
The evolution of Large Language Models (LLMs) has sparked growing concerns about jailbreak, crafted prompts that bypass safety guardrails and lead to the generation of harmful information. While recent research primarily focuses on identifying harmful content within LLM outputs, limited attention ha…