← Search

Sungee Hong

2 accepted papers

2025

A Principled Path to Fitted Distributional Evaluation

NeurIPS 2025spotlight

In reinforcement learning, distributional off-policy evaluation (OPE) focuses on estimating the return distribution of a target policy using offline data collected under a different policy. This work focuses on extending the widely used fitted Q-evaluation---developed for expectation-based reinforce…

Cited by 0SourcecodeScholar
2025

Distributional Off-policy Evaluation with Bellman Residual Minimization

AISTATS 2025poster

We study distributional off-policy evaluation (OPE), of which the goal is to learn the distribution of the return for a target policy using offline data generated by a different policy. The theoretical foundation of many existing work relies on the supremum-extended statistical distances such as sup…

Cited by 0SourcecodeScholar