IJCAI 20260 citations

BEACON: Budget-Efficient Discovery of Policy Violations in Large Language Models via Cognitive-Guided Monte Carlo Tree Search

Xinyi Huang, Jie Wang, Pengrui Xiang, Yifan Wang, Yu Fu, Jinduo Liu

Abstract

Systematic safety evaluation of large language models must uncover diverse policy violations under tight query budgets. However, most red-teaming methods optimize attack success rate and repeatedly probe a narrow set of vulnerabilities, yielding redundant failures and leaving rarer yet critical violation categories unexplored. Under fixed budgets, such inefficient exploration delays the first discovery and limits category coverage. To address these limitations, we propose the Budget-Efficient Adaptive Cognitive Offense Navigator (BEACON), a budget-aware safety testing framework that uses Cognitive-Guided Monte Carlo Tree Search to navigate the violation search space under fixed budgets. BEACON innovatively approaches safety testing as a budget-constrained failure discovery process, aiming to identify diverse safety violations as early as possible within a fixed query budget. It also provides an efficiency-oriented evaluation perspective that measures early discovery and harm category coverage under budget constraints. Experiments on standard benchmarks and frontier LLMs show that BEACON discovers failures earlier and achieves higher coverage across policy violation categories. These results underscore the value of evaluating safety testing through discovery efficiency rather than attack success rate alone. Warning: This paper contains examples of harmful language and images, and reader discretion is recommended.

Multidisciplinary Topics and Applications: Security and privacyNatural Language Processing: Language modelsNatural Language Processing: Resources and evaluationSearch: Applications
BibTeX
@inproceedings{ijcai2026_beaconbudgeteffi,
  title = {BEACON: Budget-Efficient Discovery of Policy Violations in Large Language Models via Cognitive-Guided Monte Carlo Tree Search},
  author = {Xinyi Huang and Jie Wang and Pengrui Xiang and Yifan Wang and Yu Fu and Jinduo Liu},
  booktitle = {IJCAI 2026},
  year = {2026}
}