2025
ReasonerRank: Redefining Language Model Evaluation with Ground-Truth-Free Ranking Frameworks
ACL 2025finding
Large Language Models (LLMs) are increasingly adopted across real-world applications, yet traditional evaluations rely on expensive, domain-specific ground-truth labels that are often unavailable or infeasible. We introduce a ground-truth-free evaluation framework focused on reasoning consistency an…