2023
Coordinate Ascent for Off-Policy RL with Global Convergence Guarantees
AISTATS 2023poster
We revisit the domain of off-policy policy optimization in RL from the perspective of coordinate ascent. One commonly-used approach is to leverage the off-policy policy gradient to optimize a surrogate objective – the total discounted in expectation return of the target policy with respect to the st…