← Search

Zonghao Ying

4 accepted papers

2026

AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions

CVPR 2026

The integration of vision-language models (VLMs) is driving a new generation of embodied agents capable of operating in human-centered environments. However, as deployment expands, these systems face growing safety risks, particularly when executing hazardous instructions. Current safety evaluation

Cited by 29SourceScholar
2026

Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?

ICASSP 2026poster

Jailbreak attacks on Large Language Models (LLMs) have demonstrated various successful methods whereby attackers manipulate models into generating harmful responses that they are designed to avoid. Among these, Greedy Coordinate Gradient (GCG) has emerged as a general and effective approach that opt…

Cited by 0SourcePDFScholar
2025

Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing

NeurIPS 2025poster

Large Language Models (LLMs) have demonstrated impressive performance across various tasks, yet they remain vulnerable to generating toxic content, necessitating detoxification strategies to ensure safe and responsible deployment. Test-time detoxification methods, which typically introduce static or…

Cited by 0SourceScholar
2025

Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models

EMNLP 2025

Multi-turn jailbreak attacks simulate real-world human interactions by engaging large language models (LLMs) in iterative dialogues, exposing critical safety vulnerabilities. However, existing methods often struggle to balance semantic coherence with attack effectiveness, resulting in either benign