2026
CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language Models
AAAI 2026technical
Counterfactual reasoning is widely recognized as one of the most challenging and intricate aspects of causality in artificial intelligence. In this paper, we evaluate the performance of large language models (LLMs) in counterfactual reasoning. In contrast to previous studies that primarily focus on