2025
SCORE: Systematic COnsistency and Robustness Evaluation for Large Language Models
NAACL 2025industry
Typical evaluations of Large Language Models (LLMs) report a single metric per dataset, often representing the model’s best-case performance under carefully selected settings. Unfortunately, this approach overlooks model robustness and reliability in real-world applications. For instance, simple par…