← Search

Anna Sokol

1 accepted papers

2025

BenchmarkCards: Standardized Documentation for Large Language Model Benchmarks

NeurIPS 2025poster

Large language models (LLMs) are powerful tools capable of handling diverse tasks. Comparing and selecting appropriate LLMs for specific tasks requires systematic evaluation methods, as models exhibit varying capabilities across different domains. However, finding suitable benchmarks is difficult gi…

Cited by 0SourcecodeScholar