Human-Centric Behavior-Aware Adaptive Off-Policy Selection
Ge Gao, Aishwarya Mandyam, Joy He-Yueya, Min Chi, Emma Brunskill
Abstract
In many human-centric environments, such as education and healthcare, the unobservability of human underlying states has been recognized as a key obstacle for understanding individual needs, thus hindering out ability to provide personalized decision-making policies. Several reinforcement learning (RL)-related approaches have been used to facilitate sequential decision-making in these settings, including off-policy selection (OPS), which aids in safely evaluating and selecting optimal policies offline. However, existing OPS algorithms are unsuitable when both the state is unobserved and the setting requires a personalized policy. To address this challenge, we propose a behavior-aware adaptive policy selection framework (HBO) that first captures potentially unique characteristics of the state from human behaviors, and then estimates when and how to intervene with less uncertainty in a timely manner, with bounded error. HBO is evaluated over two real-world human-centric applications, intelligent tutoring and sepsis treatments, where it significantly enhanced participants' long-term course outcomes and survival rates. Broadly, our work enables improved policy personalization in high-stakes domains where extensive evaluation is not possible.
BibTeX
@inproceedings{ijcai2026_humancentricbeha,
title = {Human-Centric Behavior-Aware Adaptive Off-Policy Selection},
author = {Ge Gao and Aishwarya Mandyam and Joy He-Yueya and Min Chi and Emma Brunskill},
booktitle = {IJCAI 2026},
year = {2026}
}