AAAI 2026technical0 citations

DarkBench+: An Extended Benchmark for Evaluating Dark Patterns in Large Language Models

Yaowen Liu, Shenjia Jing, Yufei Wei, Shoumin Zhang, Jinglu Zhang, Zhen Mei, Liangliang Yue, Jiarui Wang

Abstract

With the widespread deployment of large language models (LLMs) in human-computer interaction, dark patterns have extended from traditional visual interfaces to conversational AI systems. While existing research has confirmed the prevalence of dark patterns in LLMs, current evaluation benchmarks face critical challenges including limited classification coverage, overlooked risks specific to reasoning models, and inadequate consideration of cross-linguistic differences. To address these limitations, we propose DarkBench+, an extended benchmark for evaluating dark patterns in LLMs. We construct an expanded taxonomy containing 10 major categories and 24 subcategories, introduce an annotation workflow combining manual and automated methods, and design 2,088 bilingual test samples in Chinese and English. This benchmark is the first to develop specialized evaluation dimensions for reasoning models and systematically evaluates dark pattern behaviors across nearly 40 mainstream LLMs. Experimental results demonstrate significant manipulation risks in reasoning models

BibTeX
@inproceedings{aaai2026_darkbenchanexten,
  title = {DarkBench+: An Extended Benchmark for Evaluating Dark Patterns in Large Language Models},
  author = {Yaowen Liu and Shenjia Jing and Yufei Wei and Shoumin Zhang and Jinglu Zhang and Zhen Mei and Liangliang Yue and Jiarui Wang and Peng Zhang},
  booktitle = {AAAI 2026},
  year = {2026}
}
DarkBench+: An Extended Benchmark for Evaluating Dark Patterns in Large Language Models · AAAI 2026