2024
ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
NeurIPS 2024poster
Recent methodologies in LLM self-training mostly rely on LLM generating responses and filtering those with correct output answers as training data. This approach often yields a low-quality fine-tuning training set (e.g., incorrect plans or intermediate reasoning). In this paper, we develop a reinfor…