2026
Strategic Scaling of Test-Time Compute: A Bandit Learning Approach
ICLR 2026poster
Scaling test-time compute has emerged as an effective strategy for improving the performance of large language models. However, existing methods typically allocate compute uniformly across all queries, overlooking variation in query difficulty. To address this inefficiency, we formulate test-time co…