2025
ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio–Language Models
NeurIPS 2025poster
Recent advances in Audio-Language Models (ALMs) have significantly improved multimodal understanding capabilities. However, the introduction of the audio modality also brings new and unique vulnerability vectors. Previous studies have proposed jailbreak attacks that specifically target ALMs, reveali…