← Search

Jingyi Zheng

3 accepted papers

2026

JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models

ICLR 2026poster

Large Audio Language Models (LALMs) integrate the audio modality directly into the model, rather than converting speech into text and inputting text to Large Language Models (LLMs). While jailbreak attacks on LLMs have been extensively studied, the security of LALMs with audio modalities remains lar…

Cited by 0SourcecodeScholar
2025

CHASM: Unveiling Covert Advertisements on Chinese Social Media

NeurIPS 2025poster

Current benchmarks for evaluating large language models (LLMs) in social media moderation completely overlook a serious threat: covert advertisements, which disguise themselves as regular posts to deceive and mislead consumers into making purchases, leading to significant ethical and legal concerns.…

Cited by 0SourceScholar
2025

CL-Attack: Textual Backdoor Attacks via Cross-Lingual Triggers

AAAI 2025technical

Backdoor attacks significantly compromise the security of large language models by triggering them to output specific and controlled content. Currently, triggers for textual backdoor attacks fall into two categories: fixed-token triggers and sentence-pattern triggers. However, the former are typical…