← Search

Wonbeen Oh

2 accepted papers

2025

AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners

NeurIPS 2025poster

Self-Taught Reasoners (STaR), synonymously known as Rejection sampling Fine-Tuning (RFT), is an integral part of the training pipeline of self-improving reasoning Language Models (LMs). The self-improving mechanism often employs random observation (data) sampling. However, this results in trained…

Cited by 0SourceScholar
2025

FlickerFusion: Intra-trajectory Domain Generalizing Multi-agent Reinforcement Learning

ICLR 2025poster

Multi-agent reinforcement learning has demonstrated significant potential in addressing complex cooperative tasks across various real-world applications. However, existing MARL approaches often rely on the restrictive assumption that the number of entities (e.g., agents, obstacles) remains constant…