← Search

Gyuho Shim

1 accepted papers

2025

Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks

EMNLP 2025

Large Language Models are commonly judged by their scores on standard benchmarks, yet such scores often overstate real capability since they mask the mix of skills a task actually demands. For example, ARC is assumed to test reasoning, while HellaSwag is designed to evaluate commonsense. However, we

Cited by 0SourcePDFScholar