← Search

Wenhan Dong

2 accepted papers

2026

JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models

ICLR 2026poster

Large Audio Language Models (LALMs) integrate the audio modality directly into the model, rather than converting speech into text and inputting text to Large Language Models (LLMs). While jailbreak attacks on LLMs have been extensively studied, the security of LALMs with audio modalities remains lar…

Cited by 0SourcecodeScholar
2025

CHASM: Unveiling Covert Advertisements on Chinese Social Media

NeurIPS 2025poster

Current benchmarks for evaluating large language models (LLMs) in social media moderation completely overlook a serious threat: covert advertisements, which disguise themselves as regular posts to deceive and mislead consumers into making purchases, leading to significant ethical and legal concerns.…

Cited by 0SourceScholar