← Search

Huan Deng

2 accepted papers

2025

Refusal-Aware Red Teaming: Exposing Inconsistency in Safety Evaluations

EMNLP 2025

The responsible deployment of Large Language Models (LLMs) necessitates rigorous safety evaluations. However, a critical challenge arises from inconsistencies between an LLM’s internal refusal decisions and external safety assessments, hindering effective validation. This paper introduces the concep

Cited by 0SourcePDFScholar
2024

Suitable is the Best: Task-Oriented Knowledge Fusion in Vulnerability Detection

NeurIPS 2024poster

Deep learning technologies have demonstrated remarkable performance in vulnerability detection. Existing works primarily adopt a uniform and consistent feature learning pattern across the entire target set. While designed for general-purpose detection tasks, they lack sensitivity towards target code…

Cited by 0SourcePDFScholar