← Search

Xiaotie Deng

20 accepted papers

2025

Mechanism Design for LLM Fine-tuning with Multiple Reward Models

NeurIPS 2025poster

Fine-tuning large language models (LLMs) to aggregate multiple preferences has attracted considerable research attention. With aggregation algorithms advancing, a potential economic scenario arises where fine-tuning services are provided to agents with different preferences. In this context, agents…

Cited by 0SourceScholar
2025

Survey on Strategic Mining in Blockchain: A Reinforcement Learning Approach

IJCAI 2025

Strategic mining attacks, such as selfish mining, exploit blockchain consensus protocols by deviating from honest behavior to maximize rewards. Markov Decision Process (MDP) analysis faces scalability challenges in modern digital economics, including blockchain. To address these limitations, reinfor

Cited by 0SourcePDFScholar
2024

Contextual Decision-Making with Knapsacks Beyond the Worst Case

NeurIPS 2024poster

We study the framework of a dynamic decision-making scenario with resource constraints. In this framework, an agent, whose target is to maximize the total reward under the initial inventory, selects an action in each round upon observing a random request, leading to a reward and resource consumption…

Cited by 0SourcePDFScholar
2024

Dynamic Budget Throttling in Repeated Second-Price Auctions

AAAI 2024technical

In today's online advertising markets, a crucial requirement for an advertiser is to control her total expenditure within a time horizon under some budget. Among various budget control methods, throttling has emerged as a popular choice, managing an advertiser's total expenditure by selecting only…

Cited by 10SourcePDFScholar
2024

Learning Thresholds with Latent Values and Censored Feedback

ICLR 2024poster

In this paper, we investigate a problem of *actively* learning threshold in latent space, where the *unknown* reward $g(\gamma, v)$ depends on the proposed threshold $\gamma$ and latent value $v$ and it can be $only$ achieved if the threshold is lower than or equal to the *unknown* latent value. Thi…

Cited by 0SourcePDFScholar
2023

A Scalable Neural Network for DSIC Affine Maximizer Auction Design

NeurIPS 2023spotlight

Automated auction design aims to find empirically high-revenue mechanisms through machine learning. Existing works on multi item auction scenarios can be roughly divided into RegretNet-like and affine maximizer auctions (AMAs) approaches. However, the former cannot strictly ensure dominant strategy…

Cited by 33SourcePDFScholar
2023

Coordinated Dynamic Bidding in Repeated Second-Price Auctions with Budgets

ICML 2023poster

In online ad markets, a rising number of advertisers are employing bidding agencies to participate in ad auctions. These agencies are specialized in designing online algorithms and bidding on behalf of their clients. Typically, an agency usually has information on multiple advertisers, so she can po…

Cited by 6SourcePDFScholar
2023

From Monopoly to Competition: Optimal Contests Prevail

AAAI 2023technical

We study competition among contests in a general model that allows for an arbitrary and heterogeneous space of contest design and symmetric contestants. The goal of the contest designers is to maximize the contestants' sum of efforts. Our main result shows that optimal contests in the monopolistic s…

Cited by 6SourcePDFScholar
2023

Learning to Bid in Repeated First-Price Auctions with Budgets

ICML 2023poster

Budget management strategies in repeated auctions have received growing attention in online advertising markets. However, previous work on budget management in online bidding mainly focused on second-price auctions. The rapid shift from second-price auctions to first-price auctions for online ads in…

Cited by 20SourcePDFScholar
2022

A Context-Integrated Transformer-Based Neural Network for Auction Design

ICML 2022spotlight

One of the central problems in auction design is developing an incentive-compatible mechanism that maximizes the auctioneer’s expected revenue. While theoretical approaches have encountered bottlenecks in multi-item auctions, recently, there has been much progress on finding the optimal mechanism th…

2022

Multiagent Q-learning with Sub-Team Coordination

NeurIPS 2022accept

In many real-world cooperative multiagent reinforcement learning (MARL) tasks, teams of agents can rehearse together before deployment, but then communication constraints may force individual agents to execute independently when deployed. Centralized training and decentralized execution (CTDE) is in…

Cited by 10SourcePDFScholar
2022

On the Convergence of Fictitious Play: A Decomposition Approach

IJCAI 2022poster

Fictitious play (FP) is one of the most fundamental game-theoretical learning frameworks for computing Nash equilibrium in n-player games, which builds the foundation for modern multi-agent learning algorithms. Although FP has provable convergence guarantees on zero-sum games and potential games, ma…

Cited by 4SourcePDFScholar
2021

On the Approximation of Nash Equilibria in Sparse Win-Lose Multi-player Games

AAAI 2021technical

A polymatrix game is a multi-player game over n players, where each player chooses a pure strategy from a list of its own pure strategies. The utility of each player is a sum of payoffs it gains from the two player's game from all its neighbors, under its chosen strategy and that of its neighbor. As…

Cited by 7SourcePDFScholar
2020

A Game-Theoretic Analysis of the Empirical Revenue Maximization Algorithm with Endogenous Sampling

NeurIPS 2020poster

The Empirical Revenue Maximization (ERM) is one of the most important price learning algorithms in auction design: as the literature shows it can learn approximately optimal reserve prices for revenue-maximizing auctioneers in both repeated auctions and uniform-price auctions. However, in these app…

Cited by 5SourcePDFScholar