← Search

Johann Lee

2 accepted papers

2026

Learning from Synthetic Data Improves Multi-hop Reasoning

ICLR 2026poster

Reinforcement Learning (RL) has been shown to significantly boost reasoning capabilities of large language models (LLMs) in math, coding, and multi-hop reasoning tasks. However, RL fine-tuning requires abundant high-quality verifiable data, often obtained through human-annotated datasets and LLM-as-…

Cited by 0SourcecodeScholar
2025

PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation

ICML 2025poster

High-quality benchmarks are essential for evaluating reasoning and retrieval capabilities of large language models (LLMs). However, curating datasets for this purpose is not a permanent solution as they are prone to data leakage and inflated performance results. To address these challenges, we prop…