ICLR 2026poster0 citations

RefineBench: Evaluating Refinement Capability in Language Models

Young-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon, Yechan Hwang, Jong Myoung Kim, Graham Neubig, Sean Welleck

Abstract

Can language models (LMs) self-refine their own responses? This question is increasingly relevant as more than 10% of real-world user interactions involve refinement requests (see Appendix G). Yet prior studies have largely tested LMs on verifiable tasks such as competition math or symbolic reasoning with simplified scaffolds, whereas users often pose open-ended queries and provide varying degrees of feedback about what went wrong. The recent advent of reasoning models that exhibit self-reflection patterns in their chain-of-thought further motivates this question. To address it, we introduce RefineBench, a benchmark of 1,002 challenging problems across 11 domains paired with a checklist-based evaluation framework. We evaluate two refinement modes: (1) guided refinement, where an LM is provided natural language feedback, and (2) self-refinement, where LMs attempt to improve without guidance. In the self-refinement setting, even frontier LMs such as Gemini 2.5 Pro and GPT-4.1 achieve modest baseline scores of 31.1 and 23.4, respectively, and most models fail to consistently improve across iterations (e.g., Gemini-2.5-Pro gains only +1.8%, while DeepSeek-R1 declines by –0.2%). By contrast, in guided refinement, both proprietary LMs and large open-weight LMs (>70B) can leverage targeted feedback to refine responses to near-perfect levels within five turns. These findings suggest that frontier LMs require breakthroughs to self-refine effectively when their initial responses are incorrect, and that RefineBench provides a valuable testbed for tracking progress.

RefinementLarge Language ModelChecklist
BibTeX
@inproceedings{
lee2026refinebench,
title={RefineBench: Evaluating Refinement Capability in Language Models},
author={Young-Jun Lee and Seungone Kim and Byung-Kwan Lee and Minkyeong Moon and Yechan Hwang and Jong Myoung Kim and Graham Neubig and Sean Welleck and Ho-Jin Choi},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=GYJFJz9Dy5}
}
RefineBench: Evaluating Refinement Capability in Language Models · ICLR 2026