2025
Model Consistency as a Cheap yet Predictive Proxy for LLM Elo Scores
EMNLP 2025
New large language models (LLMs) are being released every day. Some perform significantly better or worse than expected given their parameter count. Therefore, there is a need for a method to independently evaluate models. The current best way to evaluate a model is to measure its Elo score by compa