← Search

Woosung Kim

4 accepted papers

2025

FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement Learning

NeurIPS 2025poster

Multi-objective reinforcement learning (MORL) aims to optimize policies in the presence of conflicting objectives, where linear scalarization is commonly used to reduce vector-valued returns into scalar signals. While effective for certain preferences, this approach cannot capture fairness-oriented…

Cited by 0SourceScholar
2024

ROIDICE: Offline Return on Investment Maximization for Efficient Decision Making

NeurIPS 2024poster

In this paper, we propose a novel policy optimization framework that maximizes Return on Investment (ROI) of a policy using a fixed dataset within a Markov Decision Process (MDP) equipped with a cost function. ROI, defined as the ratio between the return and the accumulated cost of a policy, serves…

Cited by 0SourcePDFScholar
2024

Relaxed Stationary Distribution Correction Estimation for Improved Offline Policy Optimization

AAAI 2024technical

One of the major challenges of offline reinforcement learning (RL) is dealing with distribution shifts that stem from the mismatch between the trained policy and the data collection policy. Stationary distribution correction estimation algorithms (DICE) have addressed this issue by regularizing the…

Cited by 1SourcePDFScholar