← Search

Yoojin Oh

2 accepted papers

2025

LPOI: Listwise Preference Optimization for Vision Language Models

ACL 2025long

Aligning large VLMs with human preferences is a challenging task, as methods like RLHF and DPO often overfit to textual information or exacerbate hallucinations. Although augmenting negative image samples partially addresses these pitfalls, no prior work has employed listwise preference optimization…

2021

Learning to Arbitrate Human and Robot Control using Disagreement between Sub-Policies

IROS 2021poster

In the context of teleoperation, arbitration refers to deciding how to blend between human and autonomous robot commands. We present a reinforcement learning solution that learns an optimal arbitration strategy that allocates more control authority to the human when the robot comes across a decision…

Cited by 11SourceScholar