2025
LMR-BENCH: Evaluating LLM Agent’s Ability on Reproducing Language Modeling Research
EMNLP 2025
Large language model (LLM) agents have demonstrated remarkable potential in advancing scientific discovery. However, their capability in the fundamental yet crucial task of reproducing code from research papers, especially in the NLP domain, remains underexplored. This task includes unique complex r