← Search

Tyler Wong

1 accepted papers

2025

InductionBench: LLMs Fail in the Simplest Complexity Class

ACL 2025long

Large language models (LLMs) have shown remarkable improvements in reasoning and many existing benchmarks have been addressed by models such as o1 and o3 either fully or partially. However, a majority of these benchmarks emphasize deductive reasoning, including mathematical and coding tasks in which…