← Search

Doo Hwan Hwang

2 accepted papers

2026

Group-Normalized Implicit Value Optimization for Language Models

ICLR 2026poster

Fine-tuning Large Language Models (LLMs) with reinforcement learning (RL) has become a key technique for enhancing performance on a wide range of tasks, from user alignment to complex reasoning. However, this approach is often hindered by the difficulty of fine-grained credit assignment, as it typic…

Cited by 0SourceScholar