2023
PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change
NeurIPS 2023poster
Generating plans of action, and reasoning about change have long been considered a core competence of intelligent agents. It is thus no surprise that evaluating the planning and reasoning capabilities of large language models (LLMs) has become a hot topic of research. Most claims about LLM planning…