← Search

Yecheng Ma

3 accepted papers

2023

Uniformly Conservative Exploration in Reinforcement Learning

AISTATS 2023poster

A key challenge to deploying reinforcement learning in practice is avoiding excessive (harmful) exploration in individual episodes. We propose a natural constraint on exploration—uniformly outperforming a conservative policy (adaptively estimated from all data observed thus far), up to a per-episode…

Cited by 6SourcePDFScholar
2022

Versatile Offline Imitation from Observations and Examples via Regularized State-Occupancy Matching

ICML 2022spotlight

We propose State Matching Offline DIstribution Correction Estimation (SMODICE), a novel and versatile regression-based offline imitation learning algorithm derived via state-occupancy matching. We show that the SMODICE objective admits a simple optimization procedure through an application of Fenche…

2021

State Relevance for Off-Policy Evaluation

ICML 2021spotlight

Importance sampling-based estimators for off-policy evaluation (OPE) are valued for their simplicity, unbiasedness, and reliance on relatively few assumptions. However, the variance of these estimators is often high, especially when trajectories are of different lengths. In this work, we introduce O…