From ``Sure" to ``Sorry": Detecting Jailbreak in Large Vision Language Model via JailNeurons
Large Vision-Language Models (LVLMs) are vulnerable to jailbreak attacks that can generate harmful content. Existing detection methods are either limited to detecting specific attack types or are too time-consuming, making them impractical for real-world deployment. To address these challenges, we p…