← Search

Simon Holk

7 accepted papers

2025

Flora: Sample-Efficient Preference-Based Rl Via Low-Rank Style Adaptation of Reward Functions

ICRA 2025

Preference-based reinforcement learning (PbRL) is a suitable approach for style adaptation of pre-trained robotic behavior: adapting the robot's policy to follow human user preferences while still being able to perform the original task. However, collecting preferences for the adaptation process in

Cited by 2SourcecodeScholar
2025

The Impact of VR and 2D Interfaces on Human Feedback in Preference-Based Robot Learning

IROS 2025

Aligning robot navigation with human preferences is essential for ensuring comfortable, and predictable robot movement in shared spaces. While preference-based learning methods, such as reinforcement learning from human feedback (RLHF), enable this alignment, the choice of the preference collection

Cited by 3SourceScholar
2024

SEQUEL: Semi-Supervised Preference-based RL with Query Synthesis via Latent Interpolation

ICRA 2024poster

Preference-based reinforcement learning (RL) poses as a recent research direction in robot learning, by allowing humans to teach robots through preferences on pairs of desired behaviours. Nonetheless, to obtain realistic robot policies, an arbitrarily large number of queries is required to be answer…

Cited by 3SourceScholar
2023

Aligning Human Preferences with Baseline Objectives in Reinforcement Learning

ICRA 2023poster

Practical implementations of deep reinforcement learning (deep RL) have been challenging due to an amplitude of factors, such as designing reward functions that cover every possible interaction. To address the heavy burden of robot reward engineering, we aim to leverage subjective human preferences…

Cited by 17SourceScholar
2023

VARIQuery: VAE Segment-Based Active Learning for Query Selection in Preference-Based Reinforcement Learning

IROS 2023poster

Human-in-the-loop reinforcement learning (RL) methods actively integrate human knowledge to create reward functions for various robotic tasks. Learning from preferences shows promise as alleviates the requirement of demonstrations by querying humans on state-action sequences. However, the limited gr…

Cited by 8SourceScholar