← Search

Noy Sternlicht

1 accepted papers

2025

Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation

EMNLP 2025

We introduce Debate Speech Evaluation as a novel and challenging benchmark for assessing LLM judges. Evaluating debate speeches requires a deep understanding of the speech at multiple levels, including argument strength and relevance, the coherence and organization of the speech, the appropriateness