2026
TurtleBench: Evaluating Top Language Models via Real-World Yes/No Puzzles
ICASSP 2026poster
As the application of Large Language Models (LLMs) expands, the demand for reliable evaluations increases. Existing LLM evaluation benchmarks primarily rely on static datasets, making it challenging to assess model performance in dynamic interactions with users. Moreover, these benchmarks often depe…