2025
LSRL: Process-Supervised GRPO on Latent Recurrent States Improves Mathematical Reasoning
EMNLP 2025
Latent-recurrent language models solve tasks by iteratively refining hidden states rather than emitting chain-of-thought tokens, yet the opacity of those hidden trajectories hinders credit assignment and limits mathematical reasoning accuracy. We propose Latent-State Supervised Reinforcement Learnin