← Search

Haoyang Ling

1 accepted papers

2024

NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes

ACL 2024long

Complex reasoning ability is one of the most important features of Large Language Models (LLMs). Numerous benchmarks have been established to assess the reasoning abilities of LLMs. However, they are inadequate in offering a rigorous evaluation and prone to the risk of overfitting, as these publicly…