← Search

Eric Yang Yu

5 accepted papers

2026

ReFORM: Reflected Flows for On-support Offline RL via Noise Manipulation

ICLR 2026poster

Offline reinforcement learning (RL) aims to learn the optimal policy from a fixed dataset generated by behavior policies without additional environment interactions. One common challenge that arises in this setting is the out-of-distribution (OOD) error, which occurs when the policy leaves the train…

Cited by 0SourcecodeScholar
2026

SEED-SET: Scalable Evolving Experimental Design for System-level Ethical Testing

ICLR 2026poster

As autonomous systems such as drones, become increasingly deployed in high-stakes, human-centric domains, it is critical to evaluate the ethical alignment since failure to do so imposes imminent danger to human lives, and long term bias in decision-making. Automated ethical benchmarking of these sys…

Cited by 0SourceScholar
2026

Solving Parameter-Robust Avoid Problems with Unknown Feasibility using Reinforcement Learning

ICLR 2026poster

Recent advances in deep reinforcement learning (RL) have achieved strong results on high-dimensional control tasks, but applying RL to reachability problems raises a fundamental mismatch: reachability seeks to maximize the set of states from which a system remains safe indefinitely, while RL optimiz…

Cited by 0SourceScholar
2025

Safe Beyond the Horizon: Efficient Sampling-based MPC with Neural Control Barrier Functions

RSS 2025poster

A common problem when using model predictive control (MPC) in practice is the satisfaction of safety beyond the prediction horizon. While theoretical works have shown that safety can be guaranteed by enforcing a suitable terminal set constraint or a sufficiently long prediction horizon, these techni…

Cited by 0PDFScholar
2022

Policy Optimization with Advantage Regularization for Long-Term Fairness in Decision Systems

NeurIPS 2022accept

Long-term fairness is an important factor of consideration in designing and deploying learning-based decision systems in high-stake decision-making contexts. Recent work has proposed the use of Markov Decision Processes (MDPs) to formulate decision-making with long-term fairness requirements in dyna…