← Search

Xiaoming Sun

10 accepted papers

2026

DR-Submodular Maximization with Stochastic Biased Gradients: Classical and Quantum Gradient Algorithms

ICLR 2026poster

In this work, we investigate DR-submodular maximization using stochastic biased gradients, which is a more realistic but challenging setting than stochastic unbiased gradients. We first generalize the Lyapunov framework to incorporate biased stochastic gradients, characterizing the adverse impacts o…

Cited by 0SourceScholar
2026

Distribution-Aware Energy Minimization: Physical-Inspired Efficient Active Learning and Quantum Potentials

IJCAI 2026

Active learning aims to maximize model performance with minimal annotation costs by selecting the most informative samples from large unlabeled pools, which often face a budget dilemma: uncertainty-based methods induce redundancy under low budgets, while representativeness-based methods struggle to

Cited by 0Scholar
2025

Quantum Speedups for Minimax Optimization and Beyond

NeurIPS 2025poster

This paper investigates convex-concave minimax optimization problems where only the function value access is allowed. We introduce a class of Hessian-aware quantum zeroth-order methods that can find the $\epsilon$-saddle point within $\tilde{\mathcal{O}}(d^{2/3}\epsilon^{-2/3})$ function value oracl…

Cited by 0SourceScholar
2023

Bandit Multi-linear DR-Submodular Maximization and Its Applications on Adversarial Submodular Bandits

ICML 2023poster

We investigate the online bandit learning of the monotone multi-linear DR-submodular functions, designing the algorithm $\mathtt{BanditMLSM}$ that attains $O(T^{2/3}\log T)$ of $(1-1/e)$-regret. Then we reduce submodular bandit with partition matroid constraint and bandit sequential monotone maximiz…

Cited by 12SourcePDFScholar
2023

Quantum Multi-Armed Bandits and Stochastic Linear Bandits Enjoy Logarithmic Regrets

AAAI 2023technical

Multi-arm bandit (MAB) and stochastic linear bandit (SLB) are important models in reinforcement learning, and it is well-known that classical algorithms for bandits with time horizon T suffer from the regret of at least the square root of T. In this paper, we study MAB and SLB with quantum reward or…

Cited by 22SourcePDFScholar
2022

Bounded Memory Adversarial Bandits with Composite Anonymous Delayed Feedback

IJCAI 2022poster

We study the adversarial bandit problem with composite anonymous delayed feedback. In this setting, losses of an action are split into d components, spreading over consecutive rounds after the action is chosen. And in each round, the algorithm observes the aggregation of losses that come from the la…

Cited by 2SourcePDFScholar
2022

Online Influence Maximization with Node-Level Feedback Using Standard Offline Oracles

AAAI 2022technical

We study the online influence maximization (OIM) problem in social networks, where in multiple rounds the learner repeatedly chooses seed nodes to generate cascades, observes the cascade feedback, and gradually learns the best seeds that generate the largest cascade. We focus on two major challenges…

Cited by 13SourcePDFScholar