← Search

Elizabeth Daly

2 accepted papers

2024

Ranking Large Language Models without Ground Truth

ACL 2024findings

Evaluation and ranking of large language models (LLMs) has become an important problem with the proliferation of these models and their impact. Evaluation methods either require human responses which are expensive to acquire or use pairs of LLMs to evaluate each other which can be unreliable. In thi…

Cited by 3SourcePDFScholar
2021

What Changed? Interpretable Model Comparison

IJCAI 2021poster

We consider the problem of distinguishing two machine learning (ML) models built for the same task in a human-interpretable way. As models can fail or succeed in different ways, classical accuracy metrics may mask crucial qualitative differences. This problem arises in a few contexts. In business ap…