← Search

Yilin Bao

1 accepted papers

2025

Offline Reinforcement Learning for LLM Multi-step Reasoning

ACL 2025finding

Improving the multi-step reasoning ability of large language models (LLMs) with offline reinforcement learning (RL) is essential for quickly adapting them to complex tasks. While Direct Preference Optimization (DPO) has shown promise in aligning LLMs with human preferences, it is less suitable for m…