← Search

James V Roggeveen

2 accepted papers

2026

CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers

ICLR 2026poster

Large language models (LLMs) have demonstrated remarkable progress in coding and mathematical problem-solving; however, evaluation on advanced research-level problems in the hard sciences remains scarce. To fill this gap, we present \cmt, a dataset of 50 original problems covering condensed matter…

Cited by 0SourceScholar
2025

HARDMath2: A Benchmark for Applied Mathematics Built by Students as Part of a Graduate Class

NeurIPS 2025poster

Large language models (LLMs) have shown remarkable progress in mathematical problem-solving, but evaluation has largely focused on problems that have exact analytical solutions or involve formal proofs, often overlooking approximation-based problems ubiquitous in applied science and engineering. To…

Cited by 0SourcecodeScholar