← Search

Jianping Pan

3 accepted papers

2023

A near-optimal high-probability swap-Regret upper bound for multi-agent bandits in unknown general-sum games

UAI 2023poster

In this paper, we study a multi-agent bandit problem in an unknown general-sum game repeated for a number of rounds (i.e., learning in a black-box game with bandit feedback), where a set of agents have no information about the underlying game structure and cannot observe each other’s actions and rew…

Cited by 4SourcePDFScholar
2020

A Unified Model for the Two-stage Offline-then-Online Resource Allocation

IJCAI 2020poster

With the popularity of the Internet, traditional offline resource allocation has evolved into a new form, called online resource allocation. It features the online arrivals of agents in the system and the real-time decision-making requirement upon the arrival of each online agent. Both offline and o…

Cited by 0SourcePDFScholar
2019

Problem-dependent Regret Bounds for Online Learning with Feedback Graphs

UAI 2019poster

This paper addresses the stochastic multi-armed bandit problem with an undirected feedback graph. We devise a UCB-based algorithm, UCB-NE, to provide a problem-dependent regret bound that depends on a clique covering. Our algorithm obtains regret which provably scales linearly with the clique coveri…

Cited by 13SourcePDFScholar