2025
SafeQuant: LLM Safety Analysis via Quantized Gradient Inspection
NAACL 2025long
Contemporary jailbreak attacks on Large Language Models (LLMs) employ sophisticated techniques with obfuscated content to bypass safety guardrails. Existing defenses either use computationally intensive LLM verification or require adversarial fine-tuning, leaving models vulnerable to advanced attack…