← Search

Hexian Ni

2 accepted papers

2026

CoRe: Combined Rewards with Vision-Language Model Feedback for Preference-Aligned Reinforcement Learning

ICML 2026poster

Reward design remains a central challenge in reinforcement learning (RL). Hand-crafted rewards are often difficult to specify and may lead to suboptimal policies, while learned rewards from preferences can suffer from inefficiency and unstable training. Inspired by the dual nature of human learning …

Cited by 0SourceScholar
2025

SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning

IROS 2025

Preference-based Reinforcement Learning (PbRL) methods provide a solution to avoid reward engineering by learning reward models based on human preferences. However, poor feedback- and sample- efficiency still remain the problems that hinder the application of PbRL. In this paper, we present a novel

Cited by 0SourcecodeScholar