← Search

Haojin Wang

2 accepted papers

2026

Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning

ICML 2026poster

Post-training of reasoning LLMs is a holistic process that typically consists of an offline SFT stage followed by an online reinforcement learning (RL) stage. However, SFT is often optimized in isolation to maximize SFT performance alone. We show that, after identical RL training, models initialized…

Cited by 0SourceScholar
2025

Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can Produce

EMNLP 2025

Autoregressive neural language models (LMs) generate a probability distribution over tokens at each time step given a prompt. In this work, we attempt to systematically understand the probability distributions that LMs can produce, showing that some distributions are significantly harder to elicit t

Cited by 0SourcePDFScholar