← Search

Athiya Deviyani

2 accepted papers

2026

Nonparametric LLM Evaluation from Preference Data

ICML 2026poster

Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards. However, many existing approaches either rely on restrictive parametric assumptions or lack valid uncertainty quantification when flexible machine learning methods are use…

Cited by 0SourceScholar
2025

Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy

NAACL 2025findings

Meta-evaluation of automatic evaluation metrics—assessing evaluation metrics themselves—is crucial for accurately benchmarking natural language processing systems and has implications for scientific inquiry, production model development, and policy enforcement. While existing approaches to metric me…