← Search

Boxiang Lyu

7 accepted papers

2026

Model-based Offline RL via Robust Value-Aware Model Learning with Implicitly Differentiable Adaptive Weighting

ICLR 2026poster

Model-based offline reinforcement learning (RL) aims to enhance offline RL with a dynamics model that facilitates policy exploration. However, model exploitation could occur due to inevitable model errors, which degrades algorithm performance. Adversarial model learning offers a theoretical framewor…

Cited by 0SourceScholar
2025

An Instrumental Value for Data Production and its Application to Data Pricing

ICML 2025poster

We develop a framework for capturing the instrumental value of data production processes, which accounts for two key factors: (a) the context of the agent’s decision-making; (b) how much data or information the buyer already possesses. We "micro-found" our data valuation function by establishing its…

Cited by 0SourcePDFScholar
2024

Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement Learning

ICML 2024spotlight

We study risk-sensitive reinforcement learning (RL), a crucial field due to its ability to enhance decision-making in scenarios where it is essential to manage uncertainty and minimize potential adverse outcomes. Particularly, our work focuses on applying the entropic risk measure to RL problems. Wh…

Cited by 2SourcePDFScholar
2023

Addressing Budget Allocation and Revenue Allocation in Data Market Environments Using an Adaptive Sampling Algorithm

ICML 2023poster

High-quality machine learning models are dependent on access to high-quality training data. When the data are not already available, it is tedious and costly to obtain them. Data markets help with identifying valuable training data: model consumers pay to train a model, the market uses that budget t…

2023

One Policy is Enough: Parallel Exploration with a Single Policy is Near-Optimal for Reward-Free Reinforcement Learning

AISTATS 2023poster

Although parallelism has been extensively used in Reinforcement Learning (RL), the quantitative effects of parallel exploration are not well understood theoretically. We study the benefits of simple parallel exploration for reward-free RL in linear Markov decision processes (MDPs) and two-player zer…

Cited by 4SourcePDFScholar
2023

Pairwise Ranking Losses of Click-Through Rates Prediction for Welfare Maximization in Ad Auctions

ICML 2023poster

We study the design of loss functions for click-through rates (CTR) to optimize (social) welfare in advertising auctions. Existing works either only focus on CTR predictions without consideration of business objectives (e.g., welfare) in auctions or assume that the distribution over the participants…

Cited by 2SourcePDFScholar
2022

Pessimism meets VCG: Learning Dynamic Mechanism Design via Offline Reinforcement Learning

ICML 2022spotlight

Dynamic mechanism design has garnered significant attention from both computer scientists and economists in recent years. By allowing agents to interact with the seller over multiple rounds, where agents’ reward functions may change with time and are state-dependent, the framework is able to model a…

Cited by 9SourcePDFScholar