2025
Uncertainty-Guided Exploration for Efficient AlphaZero Training
NeurIPS 2025poster
AlphaZero has achieved remarkable success in complex decision-making problems through self-play and neural network training. However, its self-play process remains inefficient due to limited exploration of high-uncertainty positions, the overlooked runner-up decisions in Monte Carlo Tree Search (MCT…