← Search

Jianhong Pan

2 accepted papers

2025

Semantic Representation Attack against Aligned Large Language Models

NeurIPS 2025poster

Large Language Models (LLMs) increasingly employ alignment techniques to prevent harmful outputs. Despite these safeguards, attackers can circumvent them by crafting prompts that induce LLMs to generate harmful content. Current methods typically target exact affirmative responses, suffering from lim…

Cited by 0SourceScholar
2022

GradAuto: Energy-Oriented Attack on Dynamic Neural Networks

ECCV 2022poster

"Dynamic neural networks could adapt their structures or parameters based on different inputs. By reducing the computation redundancy for certain samples, it can greatly improve the computational efficiency without compromising the accuracy. In this paper, we investigate the robustness of dynamic ne…