← Search

Bo Xue

25 accepted papers

2026

Beyond the Lower Bound: Bridging Regret Minimization and Best Arm Identification in Lexicographic Bandits

AAAI 2026technical

In multi-objective decision-making with hierarchical preferences, lexicographic bandits provide a natural framework for optimizing multiple objectives in a prioritized order. In this setting, a learner repeatedly selects arms and observes reward vectors, aiming to maximize the reward for the highest

Cited by 0SourcePDFScholar
2026

Flow of Spans: Generalizing Language Models to Dynamic Span-Vocabulary via GFlowNets

ICLR 2026poster

Standard autoregressive language models generate text token-by-token from a fixed vocabulary, inducing a *tree-structured state space* when viewing token sampling as an action, which limits flexibility and expressiveness. Recent work introduces dynamic vocabulary by sampling retrieved text spans but…

Cited by 0SourceScholar
2026

FlowRL: Matching Reward Distributions for LLM Reasoning

ICLR 2026poster

We propose FlowRL: matching the full reward distribution via flow balancing instead of solely maximizing rewards in large language model (LLM) reinforcement learning (RL). Recent advanced reasoning models adopt reward-maximizing methods (e.g., PPO and GRPO), which tend to over-optimize dominant rewa…

Cited by 0SourcecodeScholar
2026

Offline Multi-Objective Bandits: From Logged Data to Pareto-Optimal Policies

AAAI 2026technical

Offline policy learning from logged data is a critical paradigm for enabling effective decision-making without costly online exploration. However, its application has been largely confined to single-objective problems, a stark contrast to real-world scenarios where decision-making inherently involve

Cited by 0SourcePDFScholar
2025

AceParse: A Comprehensive Dataset with Diverse Structured Texts for Academic Literature Parsing

ICASSP 2025accepted

With the development of data-centric AI, the focus has shifted from model-driven approaches to improving data quality. Academic literature, as one of the crucial types, is predominantly stored in PDF formats and needs to be parsed into texts before further processing. However, parsing diverse struct…

Cited by 0SourceScholar
2025

DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models

NAACL 2025long

Recent progress in video-text retrieval has been driven largely by advancements in model architectures and training strategies. However, the representation learning capabilities of video-text retrieval models remain constrained by low-quality and limited training data annotations. To address this is…

Cited by 0SourcePDFScholar
2025

Multi-objective Linear Reinforcement Learning with Lexicographic Rewards

ICML 2025poster

Reinforcement Learning (RL) with linear transition kernels and reward functions has recently attracted growing attention due to its computational efficiency and theoretical advancements. However, prior theoretical research in RL has primarily focused on single-objective problems, resulting in limite…

Cited by 0SourcePDFScholar
2025

Multiple Trade-offs: An Improved Approach for Lexicographic Linear Bandits

AAAI 2025technical

This paper studies lexicographic online learning within the framework of multiobjective stochastic linear bandits (MOSLB), where the agent aims to simultaneously maximize multiple objectives in a hierarchical manner. Previous literature has investigated lexicographic online learning in multiobjectiv…

Cited by 0SourcePDFScholar
2025

Problem-dependent Regret for Lexicographic Multi-Armed Bandits with Adversarial Corruptions

IJCAI 2025

This paper studies lexicographic multi-armed bandits (MAB), where after selecting an arm, the agent observes a reward vector including multiple objectives, each with a different level of importance. Although previous literature has proposed the algorithm for lexicographic MAB, their algorithm suffer

Cited by 0SourcePDFScholar
2025

Safe Online Convex Optimization with Heavy-Tailed Observation Noises

AAAI 2025technical

We investigate safe online convex optimization (SOCO), where each decision must satisfy a set of unknown linear constraints. Assuming that the unknown constraints can be observed with a sub-Gaussian noise for each chosen decision, previous studies have established a high-probability regret bound of…

Cited by 0SourcePDFScholar
2025

TS-MOF: Two-Stage Multi-Objective Fine-tuning for Long-Tailed Recognition

NeurIPS 2025poster

Long-Tailed Recognition (LTR) presents a significant challenge due to extreme class imbalance, where existing methods often struggle to balance performance across head and tail classes. Directly applying multi-objective optimization (MOO) to leverage multiple LTR strategies can be complex and unstab…

Cited by 0SourceScholar
2024

Hierarchize Pareto Dominance in Multi-Objective Stochastic Linear Bandits

AAAI 2024technical

Multi-objective Stochastic Linear bandit (MOSLB) plays a critical role in the sequential decision-making paradigm, however, most existing methods focus on the Pareto dominance among different objectives without considering any priority. In this paper, we study bandit algorithms under mixed Pareto-le…

2024

Multiobjective Lipschitz Bandits under Lexicographic Ordering

AAAI 2024technical

This paper studies the multiobjective bandit problem under lexicographic ordering, wherein the learner aims to simultaneously maximize ? objectives hierarchically. The only existing algorithm for this problem considers the multi-armed bandit model, and its regret bound is O((KT)^(2/3)) under a metri…

Cited by 3SourcePDFScholar
2023

Balance Act: Mitigating Hubness in Cross-Modal Retrieval with Query and Gallery Banks

EMNLP 2023long main

In this work, we present a post-processing solution to address the hubness problem in cross-modal retrieval, a phenomenon where a small number of gallery data points are frequently retrieved, resulting in a decline in retrieval performance. We first theoretically demonstrate the necessity of incorpo…

Cited by 0SourcecodeScholar
2023

Efficient Algorithms for Generalized Linear Bandits with Heavy-tailed Rewards

NeurIPS 2023poster

This paper investigates the problem of generalized linear bandits with heavy-tailed rewards, whose $(1+\epsilon)$-th moment is bounded for some $\epsilon\in (0,1]$. Although there exist methods for generalized linear bandits, most of them focus on bounded or sub-Gaussian rewards and are not well-sui…

Cited by 4SourcePDFScholar
2023

Exploring and Verbalizing Academic Ideas by Concept Co-occurrence

ACL 2023long

Researchers usually come up with new ideas only after thoroughly comprehending vast quantities of literature. The difficulty of this procedure is exacerbated by the fact that the number of academic publications is growing exponentially. In this study, we devise a framework based on concept co-occurr…

2023

Hypervolume Maximization: A Geometric View of Pareto Set Learning

NeurIPS 2023poster

This paper presents a novel approach to multiobjective algorithms aimed at modeling the Pareto set using neural networks. Whereas previous methods mainly focused on identifying a finite number of solutions, our approach allows for the direct modeling of the entire Pareto set. Furthermore, we estab…

2021

Deep Unified Cross-Modality Hashing by Pairwise Data Alignment

IJCAI 2021poster

With the increasing amount of multimedia data, cross-modality hashing has made great progress as it achieves sub-linear search time and low memory space. However, due to the huge discrepancy between different modalities, most existing cross-modality hashing methods cannot learn unified hash codes an…

Cited by 21SourcePDFScholar
2020

Nearly Optimal Regret for Stochastic Linear Bandits with Heavy-Tailed Payoffs

IJCAI 2020poster

In this paper, we study the problem of stochastic linear bandits with finite action sets. Most of existing work assume the payoffs are bounded or sub-Gaussian, which may be violated in some scenarios such as financial markets. To settle this issue, we analyze the linear bandits with heavy-tailed pay…

Cited by 0SourcePDFScholar