AAAI 2026technical0 citations

CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language Models

Yuefei Chen, Vivek K. Singh, Jing Ma, Ruixiang Tang

Abstract

Counterfactual reasoning is widely recognized as one of the most challenging and intricate aspects of causality in artificial intelligence. In this paper, we evaluate the performance of large language models (LLMs) in counterfactual reasoning. In contrast to previous studies that primarily focus on commonsense causal reasoning, where LLMs often rely on prior knowledge for inference, we specifically assess their ability to perform counterfactual inference using a set of formal rules. To support this evaluation, we introduce a new benchmark dataset, CounterBench, comprising 1.2K counterfactual reasoning questions. The dataset is designed with varying levels of difficulty, diverse causal graph structures, distinct types of counterfactual questions, and multiple nonsensical name variants. Our experiments demonstrate that counterfactual reasoning poses a significant challenge for LLMs, with most models performing at levels comparable to random guessing. To enhance LLM

BibTeX
@inproceedings{aaai2026_counterbencheval,
  title = {CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language Models},
  author = {Yuefei Chen and Vivek K. Singh and Jing Ma and Ruixiang Tang},
  booktitle = {AAAI 2026},
  year = {2026}
}
CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language Models · AAAI 2026