← Search

Taehyun Hwang

7 accepted papers

2026

Generalized Linear Bandits with Memory

ICML 2026poster

We study generalized linear bandits with memory, a non-stationary setting in which rewards depend on past actions through a finite memory matrix. Building on prior work for linear models Clerici et al.,(2024), we show that the previously known $\tilde{\mathcal{O}}(T^{3/4})$ regret stems from a loose…

Cited by 0SourceScholar
2024

Randomized Exploration for Reinforcement Learning with Multinomial Logistic Function Approximation

NeurIPS 2024poster

We study reinforcement learning with _multinomial logistic_ (MNL) function approximation where the underlying transition probability kernel of the _Markov decision processes_ (MDPs) is parametrized by an unknown transition core with features of state and action. For the finite horizon episodic setti…

Cited by 0SourcePDFScholar
2023

Model-Based Reinforcement Learning with Multinomial Logistic Function Approximation

AAAI 2023technical

We study model-based reinforcement learning (RL) for episodic Markov decision processes (MDP) whose transition probability is parametrized by an unknown transition core with features of state and action. Despite much recent progress in analyzing algorithms in the linear MDP setting, the understandin…

Cited by 9SourcePDFScholar