← Search

Xiaoran Jin

1 accepted papers

2024

ReFT: Reasoning with Reinforced Fine-Tuning

ACL 2024long

One way to enhance the reasoning capability of Large Language Models (LLMs) is to conduct Supervised Fine-Tuning (SFT) using Chain-of-Thought (CoT) annotations. This approach does not show sufficiently strong generalization ability, however, because the training only relies on the given CoT data. In…