← Search

Rathod Darshan D

1 accepted papers

2025

SafeQuant: LLM Safety Analysis via Quantized Gradient Inspection

NAACL 2025long

Contemporary jailbreak attacks on Large Language Models (LLMs) employ sophisticated techniques with obfuscated content to bypass safety guardrails. Existing defenses either use computationally intensive LLM verification or require adversarial fine-tuning, leaving models vulnerable to advanced attack…

Cited by 0SourcePDFScholar