2025
Evaluating Compound AI Systems through Behaviors, Not Benchmarks
EMNLP 2025
Compound AI (CAI) systems, also referred to as LLM Agents, combine LLMs with retrievers and tools to enable information-seeking applications in the real-world. Thus, ensuring these systems perform reliably is critical. However, traditional evaluation using benchmark datasets and aggregate metrics of