← Search

Yunhui Jang

8 accepted papers

2025

Self-Training Large Language Models with Confident Reasoning

EMNLP 2025

Large language models (LLMs) have shown impressive performance by generating reasoning paths before final answers, but learning such a reasoning path requires costly human supervision. To address this issue, recent studies have explored self-training methods that improve reasoning capabilities using

2024

Pessimistic Backward Policy for GFlowNets

NeurIPS 2024poster

This paper studies Generative Flow Networks (GFlowNets), which learn to sample objects proportionally to a given reward function through the trajectory of state transitions. In this work, we observe that GFlowNets tend to under-exploit the high-reward objects due to training on insufficient number o…