2025
Using tournaments to calculate AUROC for zero-shot classification with LLMs
EMNLP 2025
Large language models perform surprisingly well on many zero-shot classification tasks, but are difficult to fairly compare to supervised classifiers due to the lack of a modifiable decision boundary. In this work, we propose and evaluate a method that transforms binary classification tasks into pai