← Search

Haoran Dang

1 accepted papers

2026

Temperature as a Meta-Policy: Adaptive Temperature in LLM Reinforcement Learning

ICLR 2026poster

Temperature is a crucial hyperparameter in large language models (LLMs), controlling the trade-off between exploration and exploitation during text generation. High temperatures encourage diverse but noisy outputs, while low temperatures produce focused outputs but may cause premature convergence. Y…

Cited by 0SourceScholar