← Search

Junyan Liu

9 accepted papers

2026

Revisiting the Bertrand Paradox via Equilibrium Analysis of No-regret Learners

ICML 2026poster

We study the discrete Bertrand pricing game with a non-increasing demand function. The game has $n \ge 2$ players who simultaneously choose prices from the set {$1/k, 2/k, \ldots, 1$}, where $k\in\mathbb{N}$. The player who sets the lowest price captures the entire demand; if multiple players tie fo…

Cited by 0SourceScholar
2025

Improved Regret and Contextual Linear Extension for Pandora's Box and Prophet Inequality

NeurIPS 2025poster

We study the Pandora’s Box problem in an online learning setting with semi-bandit feedback. In each round, the learner sequentially pays to open up to $n$ boxes with unknown reward distributions, observes rewards upon opening, and decides when to stop. The utility of the learner is the maximum obser…

Cited by 0SourceScholar
2025

Learning to Incentivize in Repeated Principal-Agent Problems with Adversarial Agent Arrivals

ICML 2025poster

We initiate the study of a repeated principal-agent problem over a finite horizon $T$, where a principal sequentially interacts with $K\geq 2$ types of agents arriving in an *adversarial* order. At each round, the principal strategically chooses one of the $N$ arms to incentivize for an arriving age…

Cited by 0SourcePDFScholar
2024

Uniform Last-Iterate Guarantee for Bandits and Reinforcement Learning

NeurIPS 2024poster

Existing metrics for reinforcement learning (RL) such as regret, PAC bounds, or uniform-PAC (Dann et al., 2017), typically evaluate the cumulative performance, while allowing the play of an arbitrarily bad policy at any finite time t. Such a behavior can be highly detrimental in high-stakes applicat…

Cited by 3SourcePDFScholar
2023

Improved Best-of-Both-Worlds Guarantees for Multi-Armed Bandits: FTRL with General Regularizers and Multiple Optimal Arms

NeurIPS 2023poster

We study the problem of designing adaptive multi-armed bandit algorithms that perform optimally in both the stochastic setting and the adversarial setting simultaneously (often known as a best-of-both-world guarantee). A line of recent works shows that when configured and analyzed properly, the Fol…

Cited by 21SourcePDFScholar
2023

No-Regret Online Reinforcement Learning with Adversarial Losses and Transitions

NeurIPS 2023poster

Existing online learning algorithms for adversarial Markov Decision Processes achieve $\mathcal{O}(\sqrt{T})$ regret after $T$ rounds of interactions even if the loss functions are chosen arbitrarily by an adversary, with the caveat that the transition function has to be fixed. This is because it h…

Cited by 17SourcePDFScholar
2021

Parameter Estimation for Student's t VAR Model with Missing Data

ICASSP 2021accepted

The vector autoregressive (VAR) models provide a significant tool for multivariate time series analysis. Most existing works on VAR modeling are based on the multivariate Gaussian distribution. However, heavy-tailed distributions are suggested more reasonable for capturing the real-world phenomena,…

Cited by 0SourceScholar
2018

Parameter Estimation of Heavy-Tailed Random Walk Model from Incomplete Data

ICASSP 2018accepted

This paper proposes a novel and structured framework for parameter estimation from incomplete time series data under heavy-tailed random walk model. Traditionally, maximum likelihood estimation (MLE) for Gaussian random walk model from incomplete data has been considered. However, it is not applicab…

Cited by 0SourceScholar