← Search

Laixi Shi

21 accepted papers

2026

Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in Reinforcement Learning

ICML 2026poster

Low-Rank Adaptation (LoRA) and its variants enable parameter-efficient fine-tuning of large language models under the supervised fine-tuning (SFT) paradigm. However, their efficacy and behavior under Reinforcement Learning with Verifiable Rewards (RLVR) are less well understood. In particular, two s…

Cited by 0SourceScholar
2026

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization

ICML 2026poster

Proactive large language model (LLM) agents aim to actively plan, query, and interact over multiple turns, enabling efficient task completion beyond passive instruction following and making them essential for real-world, user-centric applications. Agentic reinforcement learning (RL) has recently eme…

Cited by 0SourceScholar
2025

Breaking the Curse of Multiagency in Robust Multi-Agent Reinforcement Learning

ICML 2025poster

Standard multi-agent reinforcement learning (MARL) algorithms are vulnerable to sim-to-real gaps. To address this, distributionally robust Markov games (RMGs) have been proposed to enhance robustness in MARL by optimizing the worst-case performance when game dynamics shift within a prescribed uncert…

Cited by 5SourcePDFScholar
2025

Hybrid Transfer Reinforcement Learning: Provable Sample Efficiency from Shifted-Dynamics Data

AISTATS 2025oral

Online reinforcement learning (RL) typically requires online interaction data to learn a policy for a target task, but collecting such data can be high-stakes. This prompts interest in leveraging historical data to improve sample efficiency. The historical data may come from outdated or related sour…

Cited by 0SourcecodeScholar
2025

Overcoming the Curse of Dimensionality in Reinforcement Learning Through Approximate Factorization

ICML 2025poster

Factored Markov Decision Processes (FMDPs) offer a promising framework for overcoming the curse of dimensionality in reinforcement learning (RL) by decomposing high-dimensional MDPs into smaller and independently evolving components. Despite their potential, existing studies on FMDPs face three key…

Cited by 1SourcePDFScholar
2025

Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning

ICLR 2025poster

Driven by inherent uncertainty and the sim-to-real gap, robust reinforcement learning (RL) seeks to improve resilience against the complexity and variability in agent-environment sequential interactions. Despite the existence of a large number of RL benchmarks, there is a lack of standardized benchm…

Cited by 1SourcePDFScholar
2025

SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real Transfer

NeurIPS 2025poster

Deploying reinforcement learning (RL) safely in the real world is challenging, as policies trained in simulators must face the inevitable *sim-to-real gap*. Robust safe RL techniques are provably safe, however difficult to scale, while domain randomization is more practical yet prone to unsafe behav…

Cited by 0SourceScholar
2025

Tractable Multi-Agent Reinforcement Learning through Behavioral Economics

ICLR 2025oral

A significant roadblock to the development of principled multi-agent reinforcement learning (MARL) algorithms is the fact that desired solution concepts like Nash equilibria may be intractable to compute. We show how one can overcome this obstacle by introducing concepts from behavioral economics in…

Cited by 0SourcePDFScholar
2024

BECAUSE: Bilinear Causal Representation for Generalizable Offline Model-based Reinforcement Learning

NeurIPS 2024poster

Offline model-based reinforcement learning (MBRL) enhances data efficiency by utilizing pre-collected datasets to learn models and policies, especially in scenarios where exploration is costly or infeasible. Nevertheless, its performance often suffers from the objective mismatch between model and po…

Cited by 0SourcePDFScholar
2024

Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation

NeurIPS 2024poster

Safe reinforcement learning (RL) is crucial for deploying RL agents in real-world applications, as it aims to maximize long-term rewards while satisfying safety constraints. However, safe RL often suffers from sample inefficiency, requiring extensive interactions with the environment to learn a safe…

2024

Federated Offline Reinforcement Learning: Collaborative Single-Policy Coverage Suffices

ICML 2024poster

Offline reinforcement learning (RL), which seeks to learn an optimal policy using offline data, has garnered significant interest due to its potential in critical applications where online data collection is infeasible or expensive. This work explores the benefit of federated learning for offline RL…

Cited by 11SourcePDFScholar
2024

Near-Optimal Distributionally Robust Reinforcement Learning with General $L_p$ Norms

NeurIPS 2024poster

To address the challenges of sim-to-real gap and sample efficiency in reinforcement learning (RL), this work studies distributionally robust Markov decision processes (RMDPs) --- optimize the worst-case performance when the deployed environment is within an uncertainty set around some nominal MDP. D…

Cited by 0SourcePDFScholar
2024

Sample-Efficient Robust Multi-Agent Reinforcement Learning in the Face of Environmental Uncertainty

ICML 2024poster

To overcome the sim-to-real gap in reinforcement learning (RL), learned policies must maintain robustness against environmental uncertainties. While robust RL has been widely studied in single-agent regimes, in multi-agent environments, the problem remains understudied---despite the fact that the pr…

Cited by 13SourcePDFScholar
2023

A trajectory is worth three sentences: multimodal transformer for offline reinforcement learning

UAI 2023poster

Transformers hold tremendous promise in solving offline reinforcement learning (RL) by formulating it as a sequence modeling problem inspired by language modeling (LM). Prior works using transformers model a sample (trajectory) of RL as one sequence analogous to a sequence of words (one sentence) in…

Cited by 11SourcePDFScholar
2023

Seeing is not Believing: Robust Reinforcement Learning against Spurious Correlation

NeurIPS 2023poster

Robustness has been extensively studied in reinforcement learning (RL) to handle various forms of uncertainty such as random perturbations, rare events, and malicious attacks. In this work, we consider one critical type of robustness against spurious correlation, where different portions of the stat…

Cited by 26SourcePDFScholar
2023

The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model

NeurIPS 2023poster

This paper investigates model robustness in reinforcement learning (RL) via the framework of distributionally robust Markov decision processes (RMDPs). Despite recent efforts, the sample complexity of RMDPs is much less understood regardless of the uncertainty set in use; in particular, there exist…

Cited by 44SourcePDFScholar
2022

Curriculum Reinforcement Learning using Optimal Transport via Gradual Domain Adaptation

NeurIPS 2022accept

Curriculum Reinforcement Learning (CRL) aims to create a sequence of tasks, starting from easy ones and gradually learning towards difficult tasks. In this work, we focus on the idea of framing CRL as interpolations between a source (auxiliary) and a target task distribution. Although existing studi…

2022

Pessimistic Q-Learning for Offline Reinforcement Learning: Towards Optimal Sample Complexity

ICML 2022spotlight

Offline or batch reinforcement learning seeks to learn a near-optimal policy using history data without active exploration of the environment. To counter the insufficient coverage and sample scarcity of many offline datasets, the principle of pessimism has been recently introduced to mitigate high b…

Cited by 116SourcePDFScholar
2021

Breaking the Sample Complexity Barrier to Regret-Optimal Model-Free Reinforcement Learning

NeurIPS 2021spotlight

Achieving sample efficiency in online episodic reinforcement learning (RL) requires optimally balancing exploration and exploitation. When it comes to a finite-horizon episodic Markov decision process with $S$ states, $A$ actions and horizon length $H$, substantial progress has been achieved toward…

Cited by 64SourcePDFScholar
2021

Fusion-Based Digital Image Correlation Framework for Strain Measurement

ICASSP 2021accepted

We address the problem of enabling two-dimensional digital image correlation (DIC) for strain measurement on large three-dimensional objects with curved surfaces. It is challenging to acquire full-field qualified images of the surface required by DIC due to geometric distortion and the narrow visual…

Cited by 0SourceScholar
2020

Manifold Gradient Descent Solves Multi-Channel Sparse Blind Deconvolution Provably and Efficiently

ICASSP 2020accepted

Multi-channel sparse blind deconvolution refers to the problem of learning an unknown filter by observing its circulant convolutions with multiple input signals that are sparse. It is challenging to learn the filter efficiently due to the bilinear structure of the observations with respect to the un…

Cited by 0SourceScholar