← Search

Shaohui Mei

2 accepted papers

2025

Semantic Representation Attack against Aligned Large Language Models

NeurIPS 2025poster

Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting prompts that induce LLMs to generate harmful content. Current methods typically target exact affirmative responses, suffering from lim…

Cited by 0SourceScholar