ARBench: Algorithmic Reasoner or API Alchemist? Evaluating LLMs Beyond API Calls
Large Language Models (LLMs) have demonstrated impressive capabilities in code generation. Like human programmers, LLMs tend to call high-level APIs and libraries to program efficiently. However, this shortcut may hinder LLMs from learning the essential algorithm reasoning, leading instead to rote m