← Search

Yu Heng Hung

3 accepted papers

2026

Test-Time Alignment for Large Language Models via Textual Model Predictive Control

ICLR 2026poster

Aligning Large Language Models (LLMs) with human preferences through finetuning is resource-intensive, motivating lightweight alternatives at test time. We address test-time alignment through the lens of sequential decision making, a perspective that reveals two fundamental challenges. When actions…

Cited by 0SourceScholar
2025

BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL

ICLR 2025poster

Bayesian optimization (BO) offers an efficient pipeline for optimizing black-box functions with the help of a Gaussian process prior and an acquisition function (AF). Recently, in the context of single-objective BO, learning-based AFs witnessed promising empirical results given its favorable non-myo…

Cited by 0SourcePDFScholar
2020

Exploration Through Reward Biasing: Reward-Biased Maximum Likelihood Estimation for Stochastic Multi-Armed Bandits

ICML 2020poster

Inspired by the Reward-Biased Maximum Likelihood Estimate method of adaptive control, we propose RBMLE – a novel family of learning algorithms for stochastic multi-armed bandits (SMABs). For a broad range of SMABs including both the parametric Exponential Family as well as the non-parametric sub-Gau…

Cited by 16SourcePDFScholar