← Search

Yihan Du

16 accepted papers

2024

Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization

ICML 2024poster

Reinforcement Learning from Human Feedback (RLHF) has achieved impressive empirical successes while relying on a small amount of human feedback. However, there is limited theoretical justification for this phenomenon. Additionally, most recent studies focus on value-based algorithms despite the rece…

Cited by 16SourcePDFScholar
2024

Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human Feedback

ICLR 2024poster

Risk-sensitive reinforcement learning (RL) aims to optimize policies that balance the expected reward and risk. In this paper, we present a novel risk-sensitive RL framework that employs an Iterated Conditional Value-at-Risk (CVaR) objective under both linear and general function approximations, enr…

Cited by 3SourcePDFScholar
2023

Provably Efficient Risk-Sensitive Reinforcement Learning: Iterated CVaR and Worst Path

ICLR 2023poster

In this paper, we study a novel episodic risk-sensitive Reinforcement Learning (RL) problem, named Iterated CVaR RL, which aims to maximize the tail of the reward-to-go at each step, and focuses on tightly controlling the risk of getting into catastrophic situations at each stage. This formulation i…

Cited by 29SourcePDFScholar
2023

Provably Safe Reinforcement Learning with Step-wise Violation Constraints

NeurIPS 2023poster

We investigate a novel safe reinforcement learning problem with step-wise violation constraints. Our problem differs from existing works in that we focus on stricter step-wise violation constraints and do not assume the existence of safe actions, making our formulation more suitable for safety-criti…

Cited by 12SourcePDFScholar
2022

Branching Reinforcement Learning

ICML 2022spotlight

In this paper, we propose a novel Branching Reinforcement Learning (Branching RL) model, and investigate both Regret Minimization (RM) and Reward-Free Exploration (RFE) metrics for this model. Unlike standard RL where the trajectory of each episode is a single $H$-step path, branching RL allows an a…

Cited by 0SourcePDFScholar
2019

Direct Object Recognition Without Line-Of-Sight Using Optical Coherence

CVPR 2019poster

Visual object recognition under situations in which the direct line-of-sight is blocked, such as when it is occluded around the corner, is of practical importance in a wide range of applications. With coherent illumination, the light scattered from diffusive walls forms speckle patterns that contain…

Cited by 45PDFScholar