← Search

Qin Lu

14 accepted papers

2026

Approximation of Log-Partition Function in Policy Mirror Descent Induces Implicit Regularization for LLM Post-Training

ICML 2026poster

Policy mirror descent (PMD) provides a principled framework for reinforcement learning (RL) by iteratively solving KL-regularized policy improvement subproblems. While this approach has been adopted in training advanced LLMs such as Kimi K1.5/K2, the ideal closed-form PMD updates require reliable pa…

Cited by 0SourceScholar
2025

Ask a Strong LLM Judge when Your Reward Model is Uncertain

NeurIPS 2025poster

Reward model (RM) plays a pivotal role in reinforcement learning with human feedback (RLHF) for aligning large language models (LLMs). However, classical RMs trained on human preferences are vulnerable to reward hacking and generalize poorly to out-of-distribution (OOD) inputs. By contrast, strong…

Cited by 0SourceScholar
2025

LongLeader: A Comprehensive Leaderboard for Large Language Models in Long-context Scenarios

NAACL 2025long

Large Language Models (LLMs), exemplified by Claude and LLama, have exhibited impressive proficiency in tackling a myriad of Natural Language Processing (NLP) tasks. Yet, in pursuit of the ambitious goal of attaining Artificial General Intelligence (AGI), there remains ample room for enhancing LLM c…

Cited by 0SourcePDFScholar
2025

Online scalable Gaussian processes with conformal prediction for guaranteed coverage

ICASSP 2025accepted

The Gaussian process (GP) is a Bayesian non-parametric paradigm that is widely adopted for uncertainty quantification (UQ) in a number of safety-critical applications, including robotics, healthcare, as well as surveillance. The consistency of the resulting uncertainty values however, hinges on the…

Cited by 0SourceScholar
2025

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

NeurIPS 2025poster

Reinforcement learning from human feedback (RLHF) has become a powerful post-training paradigm for aligning large language models with human preferences. A core challenge in RLHF is constructing accurate reward signals, where the conventional Bradley-Terry reward models (BT RMs) often suffer from se…

Cited by 0SourcecodeScholar
2025

WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning

EMNLP 2025

While reinforcement learning (RL) has demonstrated remarkable success in enhancing large language models (LLMs), it has primarily focused on single-turn tasks such as solving math problems. Training effective web agents for multi-turn interactions remains challenging due to the complexity of long-ho

2023

Bayesian Optimization with Ensemble Learning Models and Adaptive Expected Improvement

ICASSP 2023accepted

Optimizing a black-box function that is expensive to evaluate emerges in a gamut of machine learning and artificial intelligence applications including drug discovery, policy optimization in robotics, and hyperparameter tuning of learning models to list a few. Bayesian optimization (BO) provides a p…

Cited by 0SourceScholar
2023

Recipes for Sequential Pre-training of Multilingual Encoder and Seq2Seq Models

ACL 2023findings

Pre-trained encoder-only and sequence-to-sequence (seq2seq) models each have advantages, however training both model types from scratch is computationally expensive. We explore recipes to improve pre-training efficiency by initializing one model from the other. (1) Extracting the encoder from a seq2…

2021

Gaussian Process Temporal-Difference Learning with Scalability and Worst-Case Performance Guarantees

ICASSP 2021accepted

Value function approximation is a crucial module for policy evaluation in reinforcement learning when the state space is large or continuous. The present paper revisits policy evaluation via temporal-difference (TD) learning from the Gaussian process (GP) perspective. Leveraging random features to a…

Cited by 0SourceScholar
2021

Graph-Adaptive Incremental Learning Using an Ensemble of Gaussian Process Experts

ICASSP 2021accepted

Graph-guided semi-supervised learning (SSL) is a major task emerging in a gamut of network science applications. However, most SSL approaches rely on deterministic similarity metrics for prediction, thus providing only point estimates of the sought function. To allow for uncertainty quantification,…

Cited by 0SourceScholar
2021

Online Unsupervised Learning Using Ensemble Gaussian Processes with Random Features

ICASSP 2021accepted

Gaussian process latent variable models (GPLVMs) are powerful, yet computationally heavy tools for nonlinear dimensionality reduction. Existing scalable variants utilize low- rank kernel matrix approximants that in essence subsample the embedding space. This work develops an efficient online approac…

Cited by 0SourceScholar
2020

Ensemble Gaussian Processes with Spectral Features for Online Interactive Learning with Scalability

AISTATS 2020poster

Combining benefits of kernels with Bayesian models, Gaussian process (GP) based approaches have well-documented merits not only in learning over a rich class of nonlinear functions, but also quantifying the associated uncertainty. While most GP approaches rely on a single preselected prior, the pres…

2020

Semi-Supervised Learning of Processes Over Multi-Relational Graphs

ICASSP 2020accepted

Semi-supervised learning (SSL) of dynamic processes over graphs is encountered in several applications of network science. Most of the existing approaches are unable to handle graphs with multiple relations, which arise in various real-world networks. This work deals with SSL of dynamic processes ov…

Cited by 0SourceScholar