← Search

Jiarui Gan

14 accepted papers

2026

Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games

ICML 2026poster

While Large Language Models (LLMs) excel in certain reasoning tasks, they struggle in multi-agent games where the final outcome depends on the joint strategies of all agents. In multi-agent games, the non-stationarity of other agents brings significant challenges on the evaluation of the reasoning p…

Cited by 0SourceScholar
2025

Contract Design Under Approximate Best Responses

ICML 2025poster

Principal-agent problems model scenarios where a principal aims at incentivizing an agent to take costly, unobservable actions through the provision of payments. Such interactions are ubiquitous in several real-world applications, ranging from blockchain to the delegation of machine learning tasks.…

Cited by 0SourcePDFScholar
2025

Stochastic Principal-Agent Problems: Computing and Learning Optimal History-Dependent Policies

NeurIPS 2025poster

We study a stochastic principal-agent model. A principal and an agent interact in a stochastic environment, each privy to observations about the state not available to the other. The principal has the power of commitment, both to elicit information from the agent and to signal her own information. T…

Cited by 0SourceScholar
2025

Strategyproof Reinforcement Learning from Human Feedback

NeurIPS 2025poster

We study Reinforcement Learning from Human Feedback (RLHF) in settings where multiple labelers may strategically misreport feedback to steer the learned policy toward their own preferences. We show that existing RLHF algorithms, including recent pluralistic methods, are not strategyproof, and that e…

Cited by 0SourceScholar
2023

Markov Decision Processes with Time-Varying Geometric Discounting

AAAI 2023technical

Canonical models of Markov decision processes (MDPs) usually consider geometric discounting based on a constant discount factor. While this standard modeling approach has led to many elegant results, some recent studies indicate the necessity of modeling time-varying discounting in certain applicati…

Cited by 2SourcePDFScholar
2023

Online Reinforcement Learning with Uncertain Episode Lengths

AAAI 2023technical

Existing episodic reinforcement algorithms assume that the length of an episode is fixed across time and known a priori. In this paper, we consider a general framework of episodic reinforcement learning when the length of each episode is drawn from a distribution. We first establish that this prob…

Cited by 7SourcePDFScholar
2022

Admissible Policy Teaching through Reward Design

AAAI 2022technical

We study reward design strategies for incentivizing a reinforcement learning agent to adopt a policy from a set of admissible policies. The goal of the reward designer is to modify the underlying reward function cost-efficiently while ensuring that any approximately optimal deterministic policy unde…

Cited by 17SourcePDFScholar
2020

Optimally Deceiving a Learning Leader in Stackelberg Games

NeurIPS 2020poster

Recent results in the ML community have revealed that learning algorithms used to compute the optimal strategy for the leader to commit to in a Stackelberg game, are susceptible to manipulation by the follower. Such a learning algorithm operates by querying the best responses or the payoffs of the f…

Cited by 21SourcePDFScholar
2019

Manipulating a Learning Defender and Ways to Counteract

NeurIPS 2019poster

In Stackelberg security games when information about the attacker's payoffs is uncertain, algorithms have been proposed to learn the optimal defender commitment by interacting with the attacker and observing their best responses. In this paper, we show that, however, these algorithms can be easily m…

Cited by 22SourcePDFScholar