← Search

Kihyuk Hong

7 accepted papers

2025

Generator-Mediated Bandits: Thompson Sampling for GenAI-Powered Adaptive Interventions

NeurIPS 2025poster

Recent advances in generative artificial intelligence (GenAI) models have enabled the generation of personalized content that adapts to up-to-date user context. While personalized decision systems are often modeled using bandit formulations, the integration of GenAI introduces new structure into oth…

Cited by 0SourceScholar
2025

Learning Infinite-Horizon Average-Reward Linear Mixture MDPs of Bounded Span

AISTATS 2025poster

This paper proposes a computationally tractable algorithm for learning infinite-horizon average-reward linear mixture Markov decision processes (MDPs) under the Bellman optimality condition. Our algorithm for linear mixture MDPs achieves a nearly minimax optimal regret upper bound of $\widetilde{\ma…

Cited by 0SourceScholar
2025

Reinforcement Learning for Infinite-Horizon Average-Reward Linear MDPs via Approximation by Discounted-Reward MDPs

AISTATS 2025poster

We study the problem of infinite-horizon average-reward reinforcement learning with linear Markov decision processes (MDPs). The associated Bellman operator of the problem not being a contraction makes the algorithm design challenging. Previous approaches either suffer from computational inefficienc…

Cited by 0SourceScholar
2024

A Primal-Dual Algorithm for Offline Constrained Reinforcement Learning with Linear MDPs

ICML 2024poster

We study offline reinforcement learning (RL) with linear MDPs under the infinite-horizon discounted setting which aims to learn a policy that maximizes the expected discounted cumulative reward using a pre-collected dataset. Existing algorithms for this setting either require a uniform data coverage…

Cited by 1SourcePDFScholar
2024

A Primal-Dual-Critic Algorithm for Offline Constrained Reinforcement Learning

AISTATS 2024poster

Offline constrained reinforcement learning (RL) aims to learn a policy that maximizes the expected cumulative reward subject to constraints on expected cumulative cost using an existing dataset. In this paper, we propose Primal-Dual-Critic Algorithm (PDCA), a novel algorithm for offline constrained…

Cited by 12SourcePDFScholar
2023

An Optimization-based Algorithm for Non-stationary Kernel Bandits without Prior Knowledge

AISTATS 2023poster

We propose an algorithm for non-stationary kernel bandits that does not require prior knowledge of the degree of non-stationarity. The algorithm follows randomized strategies obtained by solving optimization problems that balance exploration and exploitation. It adapts to non-stationarity by restart…

Cited by 12SourcePDFScholar