← Search

Hoang Minh Le

2 accepted papers

2022

Policy Optimization with Linear Temporal Logic Constraints

NeurIPS 2022accept

We study the problem of policy optimization (PO) with linear temporal logic (LTL) constraints. The language of LTL allows flexible description of tasks that may be unnatural to encode as a scalar cost function. We consider LTL-constrained PO as a systematic framework, decoupling task specification f…

Cited by 24SourcePDFScholar
2021

Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning

NeurIPS 2021poster

We offer an experimental benchmark and empirical study for off-policy policy evaluation (OPE) in reinforcement learning, which is a key problem in many safety critical applications. Given the increasing interest in deploying learning-based methods, there has been a flurry of recent proposals for OPE…

Cited by 176SourceScholar