← Search

Pranav Vishnu Raja

1 accepted papers

2024

A Careful Examination of Large Language Model Performance on Grade School Arithmetic

NeurIPS 2024spotlight

Large language models (LLMs) have achieved impressive success on many benchmarks for mathematical reasoning. However, there is growing concern that some of this performance actually reflects dataset contamination, where data closely resembling benchmark questions leaks into the training data, instea…

Cited by 77SourcePDFScholar