2022
Actor-Critic Policy Optimization in a Large-Scale Imperfect-Information Game
ICLR 2022poster
The deep policy gradient method has demonstrated promising results in many large-scale games, where the agent learns purely from its own experience. Yet, policy gradient methods with self-play suffer convergence problems to a Nash Equilibrium (NE) in multi-agent situations. Counterfactual regret min…