AAAI 2026technical0 citations

What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles

Mengtao Zhou, Sifan Wu, Huan Zhang, Qi Sima, Bang Liu

Abstract

We investigate the capacity of Large Language Models (LLMs) for imaginative reasoning—the proactive construction, testing, and revision of hypotheses in information-sparse environments. Existing benchmarks, often static or focused on social deduction, fail to capture the dynamic, exploratory nature of this reasoning process. To address this gap, we introduce a comprehensive research framework based on the classic "Turtle Soup" game, integrating a benchmark, an agent, and an evaluation protocol. We present TurtleSoup-Bench, the first large-scale, bilingual, interactive benchmark for imaginative reasoning, comprising 800 turtle soup stories sourced from both the Internet and expert authors. We also propose Mosaic-Agent, a novel agent designed to assess LLMs

BibTeX
@inproceedings{aaai2026_whattoasknextpro,
  title = {What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles},
  author = {Mengtao Zhou and Sifan Wu and Huan Zhang and Qi Sima and Bang Liu},
  booktitle = {AAAI 2026},
  year = {2026}
}