← Search

Henrike Beyer

2 accepted papers

2025

Lexical Recall or Logical Reasoning: Probing the Limits of Reasoning Abilities in Large Language Models

ACL 2025long

Despite the increasing interest in the reasoning abilities of Large Language Models (LLMs), existing work shows limitations in assessing logic abilities independently from lexical memory. We address this gap with Mystery-Zebra. This robust two-part benchmark (4,290 puzzles) challenges the logic abst…

2025

Natural Language Reasoning in Large Language Models: Analysis and Evaluation

ACL 2025finding

While Large Language Models (LLMs) have demonstrated promising results on a range of reasoning benchmarks—particularly in formal logic, mathematical tasks, and Chain-of-Thought prompting—less is known about their capabilities in unconstrained natural language reasoning. Argumentative reasoning, a fo…