2024
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
NeurIPS 2024poster
Despite the abundance of datasets available for assessing large language models (LLMs), the scarcity of continuous and reliable difficulty labels for individual data points, in most cases, curtails their capacity to benchmark model generalization performance across different levels of complexity. Ad…