← Search

Grigor Nalbandyan

1 accepted papers

2025

SCORE: Systematic COnsistency and Robustness Evaluation for Large Language Models

NAACL 2025industry

Typical evaluations of Large Language Models (LLMs) report a single metric per dataset, often representing the model’s best-case performance under carefully selected settings. Unfortunately, this approach overlooks model robustness and reliability in real-world applications. For instance, simple par…