← Search

Shengbo Wang

9 accepted papers

2026

Preference Is More than Comparisons: Rethinking Dueling Bandits with Augmented Human Feedback

AAAI 2026technical

Interactive preference elicitation (IPE) aims to substantially reduce human effort while acquiring human preferences in wide personalization systems. Dueling bandit (DB) algorithms enable optimal decision-making in IPE building on pairwise comparisons. However, they remain inefficient when human fee

Cited by 0SourcePDFScholar
2025

Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning

NeurIPS 2025poster

Motivated by practical applications where stable long-term performance is critical—such as robotics, operations research, and healthcare—we study the problem of distributionally robust (DR) average-reward reinforcement learning. We propose two algorithms that achieve near-optimal sample complexity.…

Cited by 0SourceScholar
2025

Statistical Learning of Distributionally Robust Stochastic Control in Continuous State Spaces

AISTATS 2025oral

We explore the control of stochastic systems with potentially continuous state and action spaces, characterized by the state dynamics $X_{t+1} = f(X_t, A_t, W_t)$. Here, $X$, $A$, and $W$ represent the state, action, and exogenous random noise processes, respectively, with $f$ denoting a known funct…

Cited by 0SourceScholar
2024

An Efficient High-dimensional Gradient Estimator for Stochastic Differential Equations

NeurIPS 2024poster

Overparameterized stochastic differential equation (SDE) models have achieved remarkable success in various complex environments, such as PDE-constrained optimization, stochastic control and reinforcement learning, financial engineering, and neural SDEs. These models often feature system evolution c…

Cited by 2SourcePDFScholar
2024

Constrained Bayesian Optimization under Partial Observations: Balanced Improvements and Provable Convergence

AAAI 2024technical

The partially observable constrained optimization problems (POCOPs) impede data-driven optimization techniques since an infeasible solution of POCOPs can provide little information about the objective as well as the constraints. We endeavor to design an efficient and provable method for expensive PO…

2024

Direct Preference-Based Evolutionary Multi-Objective Optimization with Dueling Bandits

NeurIPS 2024poster

The ultimate goal of multi-objective optimization (MO) is to assist human decision-makers (DMs) in identifying solutions of interest (SOI) that optimally reconcile multiple objectives according to their preferences. Preference-based evolutionary MO (PBEMO) has emerged as a promising framework that p…

Cited by 3SourcePDFScholar
2023

A Finite Sample Complexity Bound for Distributionally Robust Q-learning

AISTATS 2023poster

We consider a reinforcement learning setting in which the deployment environment is different from the training environment. Applying a robust Markov decision processes formulation, we extend the distributionally robust Q-learning framework studied in [Liu et. al. 2022]. Further, we improve the desi…

Cited by 36SourcePDFScholar