← Search

Giseung Park

6 accepted papers

2026

Constrained Multi-Objective Reinforcement Learning with Max-Min Criterion

ICML 2026poster

Multi-Objective Reinforcement Learning (MORL) extends standard RL by optimizing policies with respect to multiple, often conflicting, objectives. While max-min MORL has emerged as an effective approach for promoting fairness, its applicability remains limited, particularly when constraints must be i…

Cited by 0SourceScholar
2025

Multi-Objective Reinforcement Learning with Max-Min Criterion: A Game-Theoretic Approach

NeurIPS 2025poster

In this paper, we propose a provably convergent and practical framework for multi-objective reinforcement learning with max-min criterion. From a game-theoretic perspective, we reformulate max-min multi-objective reinforcement learning as a two-player zero-sum regularized continuous game and introdu…

Cited by 0SourceScholar
2024

The Max-Min Formulation of Multi-Objective Reinforcement Learning: From Theory to a Model-Free Algorithm

ICML 2024poster

In this paper, we consider multi-objective reinforcement learning, which arises in many real-world problems with multiple optimization goals. We approach the problem with a max-min framework focusing on fairness among the multiple goals and develop a relevant theory and a practical model-free algori…

2022

Blockwise Sequential Model Learning for Partially Observable Reinforcement Learning

AAAI 2022technical

This paper proposes a new sequential model learning architecture to solve partially observable Markov decision problems. Rather than compressing sequential information at every timestep as in conventional recurrent neural network-based methods, the proposed architecture generates a latent variable i…

2020

Population-Guided Parallel Policy Search for Reinforcement Learning

ICLR 2020poster

In this paper, a new population-guided parallel learning scheme is proposed to enhance the performance of off-policy reinforcement learning (RL). In the proposed scheme, multiple identical learners with their own value-functions and policies share a common experience replay buffer, and search a good…

Cited by 50SourcecodeScholar