← Search

Dohyeong Kim

18 accepted papers

2026

Compositional Transduction with Latent Analogies for Offline Goal-Conditioned Reinforcement Learning

ICML 2026poster

In offline goal-conditioned reinforcement learning (GCRL), where one relies on a limited reward-free dataset to learn a generalist goal-reaching agent, compositional generalization becomes essential for reaching unseen goals under novel contextual variations. Most prior approaches pursue this via tr…

Cited by 0SourceScholar
2026

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning

ICML 2026poster

We propose Flow-Anchored Noise-conditioned Q-Learning (FAN), a highly efficient and high-performing offline reinforcement learning (RL) algorithm. Recent work has shown that expressive flow policies and distributional critics improve offline RL performance, but at a high computational cost. Specific…

Cited by 0SourceScholar
2025

Bellman Unbiasedness: Toward Provably Efficient Distributional Reinforcement Learning with General Value Function Approximation

ICML 2025poster

Distributional reinforcement learning improves performance by capturing environmental stochasticity, but a comprehensive theoretical understanding of its effectiveness remains elusive. In addition, the intractable element of the infinite dimensionality of distributions has been overlooked. In this p…

Cited by 0SourcePDFScholar
2025

Conflict-Averse Gradient Aggregation for Constrained Multi-Objective Reinforcement Learning

ICLR 2025poster

In real-world applications, a reinforcement learning (RL) agent should consider multiple objectives and adhere to safety guidelines. To address these considerations, we propose a constrained multi-objective RL algorithm named constrained multi-objective gradient aggregator (CoMOGA). In the field of…

Cited by 0SourcePDFScholar
2025

Pareto Optimal Risk-Agnostic Distributional Bandits with Heavy-Tail Rewards

NeurIPS 2025poster

This paper addresses the problem of multi-risk measure agnostic multi-armed bandits in heavy-tailed reward settings. We propose a framework that leverages novel deviation inequalities for the $1$-Wasserstein distance to construct confidence intervals for Lipschitz risk measures. The distributional…

Cited by 0SourceScholar
2025

Policy-labeled Preference Learning: Is Preference Enough for RLHF?

ICML 2025spotlight

To design reward that align with human goals, Reinforcement Learning from Human Feedback (RLHF) has emerged as a prominent technique for learning reward functions from human preferences and optimizing models using reinforcement learning algorithms. However, existing RLHF methods often misinterpret…

Cited by 0SourcePDFScholar
2025

Stage-Wise Reward Shaping for Acrobatic Robots: A Constrained Multi-Objective Reinforcement Learning Approach

ICRA 2025

As the complexity of tasks addressed through reinforcement learning (RL) increases, the definition of reward functions also has become highly complicated. We introduce an RL method aimed at simplifying the reward-shaping process through intuitive strategies. Initially, instead of a single reward fun

Cited by 16SourcecodeScholar
2024

Adversarial Environment Design via Regret-Guided Diffusion Models

NeurIPS 2024spotlight

Training agents that are robust to environmental changes remains a significant challenge in deep reinforcement learning (RL). Unsupervised environment design (UED) has recently emerged to address this issue by generating a set of training environments tailored to the agent's capabilities. While prio…

Cited by 0SourcePDFScholar
2024

Spectral-Risk Safe Reinforcement Learning with Convergence Guarantees

NeurIPS 2024poster

The field of risk-constrained reinforcement learning (RCRL) has been developed to effectively reduce the likelihood of worst-case scenarios by explicitly handling risk-measure-based constraints. However, the nonlinearity of risk measures makes it challenging to achieve convergence and optimality. To…

2023

Dual Variable Actor-Critic for Adaptive Safe Reinforcement Learning

IROS 2023poster

Satisfying safety constraints in reinforcement learning (RL) is an important issue, especially in real-world applications. Many studies have approached safe RL with the Lagrangian method, which introduces dual variables. However, applying a trained policy with the optimal dual variable to a new envi…

Cited by 1SourceScholar
2023

SDF-Based Graph Convolutional Q-Networks for Rearrangement of Multiple Objects

ICRA 2023poster

In this paper, we propose a signed distance field (SDF)-based deep Q-learning framework for multi-object re-arrangement. Our method learns to rearrange objects with non-prehensile manipulation, e.g., pushing, in unstructured environments. To reliably estimate Q-values in various scenes, we train the…

Cited by 3SourceScholar
2023

Trust Region-Based Safe Distributional Reinforcement Learning for Multiple Constraints

NeurIPS 2023poster

In safety-critical robotic tasks, potential failures must be reduced, and multiple constraints must be met, such as avoiding collisions, limiting energy consumption, and maintaining balance. Thus, applying safe reinforcement learning (RL) in such robotic tasks requires to handle multiple constraints…

2020

MixGAIL: Autonomous Driving Using Demonstrations with Mixed Qualities

IROS 2020poster

In this paper, we consider autonomous driving of a vehicle using imitation learning. Generative adversarial imitation learning (GAIL) is a widely used algorithm for imitation learning. This algorithm leverages positive demonstrations to imitate the behavior of an expert. In this paper, we propose a…

Cited by 25SourceScholar