2026
Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions
ICASSP 2026poster
Large language models (LLMs) demonstrate remarkable performance across various tasks, prompting researchers to develop diverse evaluation benchmarks. However, most benchmarks typically measure the ability of LLMs to respond to individual questions, neglecting the complex interactions in real-world a…