2025
Lexical Recall or Logical Reasoning: Probing the Limits of Reasoning Abilities in Large Language Models
ACL 2025long
Despite the increasing interest in the reasoning abilities of Large Language Models (LLMs), existing work shows limitations in assessing logic abilities independently from lexical memory. We address this gap with Mystery-Zebra. This robust two-part benchmark (4,290 puzzles) challenges the logic abst…