Reexamining the Exploration–Exploitation Dilemma from an Entropy-Driven Perspective
Renye Yan, Yaozhong Gan, Jikang Cheng, Yi Sun, Zongwei Wang, Ling Liang, Yimao Cai
Abstract
Achieving an optimal balance between exploration and exploitation remains a fundamental challenge in reinforcement learning. This work revisits the exploration-exploitation dilemma through the lens of entropy, offering a novel perspective on this enduring problem. It establishes a theoretical connection between policy's entropy and exploratory behavior, using entropy as a rational measure to quantify the exploration-exploitation trade-off. Theoretical analyses demonstrate that a modified Bellman equation, augmented with a novelty-seeking term, ensures appropriate entropy adjustment and guarantees its globally monotonic decay. The derived policy optimization process inherently accommodates all three regimes of the exploration-exploitation spectrum, enabling a principled transition from exploration to exploitation. Building on these theoretical insights, this work introduces AdaZero, an adaptive deep architecture that dynamically balances exploration and exploitation. Extensive empirical evaluations highlight AdaZero's robust performance and validate the theory's feasibility.
BibTeX
@inproceedings{ijcai2026_reexaminingtheex,
title = {Reexamining the Exploration–Exploitation Dilemma from an Entropy-Driven Perspective},
author = {Renye Yan and Yaozhong Gan and Jikang Cheng and Yi Sun and Zongwei Wang and Ling Liang and Yimao Cai},
booktitle = {IJCAI 2026},
year = {2026}
}