← Search

Anirudhan Badrinath

2 accepted papers

2024

OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators

NeurIPS 2024poster

Offline policy evaluation (OPE) allows us to evaluate and estimate a new sequential decision-making policy's performance by leveraging historical interaction data collected from other policies. Evaluating a new policy online without a confident estimate of its performance can lead to costly, unsafe,…

Cited by 0SourcePDFScholar
2023

Waypoint Transformer: Reinforcement Learning via Supervised Learning with Intermediate Targets

NeurIPS 2023poster

Despite the recent advancements in offline reinforcement learning via supervised learning (RvS) and the success of the decision transformer (DT) architecture in various domains, DTs have fallen short in several challenging benchmarks. The root cause of this underperformance lies in their inability t…

Cited by 21SourcePDFScholar