2026
Look-ahead Reasoning with a Learned Model in Imperfect Information Games
ICLR 2026poster
Test-time reasoning significantly enhances pre-trained AI agents’ performance. However, it requires an explicit environment model, often unavailable or overly complex in real-world scenarios. While MuZero enables effective model learning for search in perfect information games, extending this paradi…