2025
BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks
NeurIPS 2025poster
Large language models (LLMs) are powerful tools capable of handling diverse tasks. Comparing and selecting appropriate LLMs for specific tasks requires systematic evaluation methods, as models exhibit varying capabilities across different domains. However, finding suitable benchmarks is difficult gi…