← Search

Eun-Ah Kim

2 accepted papers

2026

CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers

ICLR 2026poster

Large language models (LLMs) have demonstrated remarkable progress in coding and mathematical problem-solving; however, evaluation on advanced research-level problems in the hard sciences remains scarce. To fill this gap, we present \cmt, a dataset of 50 original problems covering condensed matter…

Cited by 0SourceScholar
2025

CURIE: Evaluating LLMs on Multitask Scientific Long-Context Understanding and Reasoning

ICLR 2025poster

Scientific problem-solving involves synthesizing information while applying expert knowledge. We introduce CURIE, a scientific long-Context Understanding, Reasoning, and Information Extraction benchmark to measure the potential of Large Language Models (LLMs) in scientific problem-solving a…