2024
Efficient Model-Based Concave Utility Reinforcement Learning through Greedy Mirror Descent
AISTATS 2024poster
Many machine learning tasks can be solved by minimizing a convex function of an occupancy measure over the policies that generate them. These include reinforcement learning, imitation learning, among others. This more general paradigm is called the Concave Utility Reinforcement Learning problem (CUR…