Variance Driven Exploration: A Provable and Efficient Methodology for Pure Exploration in Highly Stochastic Environments
Khang Luong, Nam Nguyen, Hoang Ta, Hung Tran-The, Tuan Dam
Abstract
We propose ***Var**iance **D**riven **E**xploration* (VarDE), a principled approach for pure exploration in *highly stochastic environments*, where the exploration process is dominated by stochastic variance. VarDE is built on a fundamental principle: *sampling effort should be allocated to minimize the uncertainty of the final decision*. We formalize the uncertainty of the final decision through a smooth decision function and derive allocation rules that explicitly capture how stochastic noise in individual components affects the reliability of the final output. We apply this methodology to three core problems of pure exploration -- Best Arm Identification (BAI), Monte Carlo Tree Search (MCTS), and Best-Policy Identification (BPI) -- with theoretical guarantees on variance decay and simple regret. Empirically, we demonstrate consistent and significant improvements of VarDE over existing methods, with especially strong gains in highly stochastic environments.
BibTeX
@inproceedings{
luong2026variance,
title={Variance Driven Exploration: A Provable and Efficient Methodology for Pure Exploration in Highly Stochastic Environments},
author={Khang Luong and Nam Nguyen and Hoang Ta and Hung The Tran and Tuan Quang Dam},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=34cmbsgpT5}
}