2026
Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
ICLR 2026poster
Recent advancements in Large Language Models (LLMs) have shifted from explicit Chain-of-Thought (CoT) reasoning to more efficient latent reasoning, where intermediate thoughts are represented as vectors rather than text. However, latent reasoning can be brittle on challenging, out-of-distribution ta…