← Search

Kazuki Ota

2 accepted papers

2026

Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games

ICML 2026poster

Two-player games such as board games have long been used as traditional benchmark for reinforcement learning. This work revisits a regularized policy optimization with reverse Kullback-Leibler divergence and entropy divergence and analyzes this combination on two-player zero-sum settings from theore…

Cited by 0SourceScholar
2025

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning

ICML 2025poster

For continuous action spaces, actor-critic methods are widely used in online reinforcement learning (RL). However, unlike RL algorithms for discrete actions, which generally model the optimal value function using the Bellman optimality operator, RL algorithms for continuous actions typically model Q…