← Search

Pengwei Zhan

7 accepted papers

2024

Rethinking Word-level Adversarial Attack: The Trade-off between Efficiency, Effectiveness, and Imperceptibility

COLING 2024main

Neural language models have demonstrated impressive performance in various tasks but remain vulnerable to word-level adversarial attacks. Word-level adversarial attacks can be formulated as a combinatorial optimization problem, and thus, an attack method can be decomposed into search space and searc…

Cited by 3SourcePDFScholar
2024

Unveiling the Lexical Sensitivity of LLMs: Combinatorial Optimization for Prompt Enhancement

EMNLP 2024main

Large language models (LLMs) demonstrate exceptional instruct-following ability to complete various downstream tasks. Although this impressive ability makes LLMs flexible task solvers, their performance in solving tasks also heavily relies on instructions. In this paper, we reveal that LLMs are over…

Cited by 4SourcePDFScholar
2023

Contrastive Learning with Adversarial Examples for Alleviating Pathology of Language Model

ACL 2023long

Neural language models have achieved superior performance. However, these models also suffer from the pathology of overconfidence in the out-of-distribution examples, potentially making the model difficult to interpret and making the interpretation methods fail to provide faithful attributions. In t…

Cited by 4SourcePDFScholar
2023

Similarizing the Influence of Words with Contrastive Learning to Defend Word-level Adversarial Text Attack

ACL 2023findings

Neural language models are vulnerable to word-level adversarial text attacks, which generate adversarial examples by directly substituting discrete input words. Previous search methods for word-level attacks assume that the information in the important words is more influential on prediction than un…

Cited by 7SourcePDFScholar
2022

Mitigating the Inconsistency Between Word Saliency and Model Confidence with Pathological Contrastive Training

ACL 2022findings

Neural networks are widely used in various NLP tasks for their remarkable performance. However, the complexity makes them difficult to interpret, i.e., they are not guaranteed right for the right reason. Besides the complexity, we reveal that the model pathology - the inconsistency between word sali…

Cited by 6SourcePDFScholar
2022

PARSE: An Efficient Search Method for Black-box Adversarial Text Attacks

COLING 2022main

Neural networks are vulnerable to adversarial examples. The adversary can successfully attack a model even without knowing model architecture and parameters, i.e., under a black-box scenario. Previous works on word-level attacks widely use word importance ranking (WIR) methods and complex search met…

Cited by 9SourcePDFScholar
2022

SP Attack: Single-Perspective Attack for Generating Adversarial Omnidirectional Images

ICASSP 2022accepted

The safety of Deep Neural Networks (DNNs) processing omnidirectional images (ODIs) is an under-researched topic. In this paper, we propose a novel sparse attack, named Single-Perspective (SP) Attack, towards fooling these models by perturbing only one perspective image (PI) rendered from the target…

Cited by 0SourceScholar