← Search

Kosuke Kawakami

2 accepted papers

2025

A General Framework for Off-Policy Learning with Partially-Observed Reward

ICLR 2025poster

Off-policy learning (OPL) in contextual bandits aims to learn a decision-making policy that maximizes the target rewards by using only historical interaction data collected under previously developed policies. Unfortunately, when rewards are only partially observed, the effectiveness of OPL degrades…

Cited by 0SourcePDFScholar
2024

Towards Assessing and Benchmarking Risk-Return Tradeoff of Off-Policy Evaluation

ICLR 2024poster

**Off-Policy Evaluation (OPE)** aims to assess the effectiveness of counterfactual policies using offline logged data and is frequently utilized to identify the top-$k$ promising policies for deployment in online A/B tests. Existing evaluation metrics for OPE estimators primarily focus on the "accur…