← Search

Guanlin Liu

6 accepted papers

2025

Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback

NeurIPS 2025poster

Reinforcement Learning from Human Feedback (RLHF) is essential for aligning large language models (LLMs) with human preferences and values. While recent research has primarily focused on algorithmic advancements—such as reducing computational overhead or strengthening reward models to mitigate rewar…

Cited by 0SourceScholar
2025

Flaming-hot Initiation with Regular Execution Sampling for Large Language Models

NAACL 2025findings

Since the release of ChatGPT, large language models (LLMs) have demonstrated remarkable capabilities across various domains. A key challenge in developing these general capabilities is efficiently sourcing diverse, high-quality data. This becomes especially critical in reasoning-related tasks with s…

Cited by 2SourcePDFScholar