2025
GroundCocoa: A Benchmark for Evaluating Compositional & Conditional Reasoning in Language Models
NAACL 2025long
The rapid progress of large language models (LLMs) has seen them excel and frequently surpass human performance on standard benchmarks. This has enabled many downstream applications, such as LLM agents, to rely on their reasoning to address complex task requirements. However, LLMs are known to unexp…