← Search

Nick Masiewicki

1 accepted papers

2025

Chatbot Arena Estimate: towards a generalized performance benchmark for LLM capabilities

NAACL 2025industry

In industrial LLM development, evaluating large language models (LLMs) is critical for tasks like benchmarking internal models and detecting regressions during fine-tuning, but existing benchmark aggregation methods, such as Elo-based systems, can be resource-intensive, public facing, and time-consu…