PRECISE: Reducing the Bias of LLM Evaluations Using Prediction-Powered Ranking Estimation
Evaluating the quality of search systems traditionally requires a significant number of human relevance annotations. In recent times, several systems have explored the usage of Large Language Models (LLMs) as automated judges for this task while their inherent biases prevent direct use for metric es