← Search

David Friede

2 accepted papers

2024

AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents

NAACL 2024system demonstrations

The advances made by Large Language Models (LLMs) have led to the pursuit of LLM agents that can solve intricate, multi-step reasoning tasks. As with any research pursuit, benchmarking and evaluation are key corner stones to efficient and reliable progress. However, existing benchmarks are often nar…