← Search

Shibei Zhu

3 accepted papers

2024

Interactive Reward Tuning: Interactive Visualization for Preference Elicitation

IROS 2024poster

In reinforcement learning, tuning reward weights in the reward function is necessary to align behavior with user preferences. However, current approaches, which use pairwise comparisons for preference elicitation, are inefficient, because they miss much of the human ability to explore and judge grou…

Cited by 1SourceScholar
2024

Preference Learning of Latent Decision Utilities with a Human-like Model of Preferential Choice

NeurIPS 2024poster

Preference learning methods make use of models of human choice in order to infer the latent utilities that underlie human behavior. However, accurate modeling of human choice behavior is challenging due to a range of context effects that arise from how humans contrast and evaluate options. Cognitive…

Cited by 0SourcePDFScholar
2023

Imitation-Guided Multimodal Policy Generation from Behaviourally Diverse Demonstrations

IROS 2023poster

Learning policies from multiple demonstrators is often difficult because different individuals perform the same task differently due to hidden factors such as preferences. In the context of policy learning, this leads to multimodal policies. Existing policy learning methods often converge to a singl…

Cited by 0SourceScholar