← Search

Anna Winnicki

3 accepted papers

2024

Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization

ICML 2024poster

Reinforcement Learning from Human Feedback (RLHF) has achieved impressive empirical successes while relying on a small amount of human feedback. However, there is limited theoretical justification for this phenomenon. Additionally, most recent studies focus on value-based algorithms despite the rece…

Cited by 16SourcePDFScholar
2023

On The Convergence Of Policy Iteration-Based Reinforcement Learning With Monte Carlo Policy Evaluation

AISTATS 2023poster

A common technique in reinforcement learning is to evaluate the value function from Monte Carlo simulations of a given policy, and use the estimated value function to obtain a new policy which is greedy with respect to the estimated value function. A well-known longstanding open problem in this cont…

Cited by 13SourcePDFScholar