← Search

Haifeng Xu

49 accepted papers

2026

Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information

ICML 2026poster

With the rapid progress of multi-agent large language model (LLM) reasoning, how to effectively aggregate answers from multiple LLMs has emerged as a fundamental challenge. Standard majority voting treats all answers equally, failing to consider latent heterogeneity and correlation across models. In…

Cited by 0SourceScholar
2026

Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence

ICLR 2026poster

Synthetic data has been increasingly used to train frontier generative models. However, recent study raises key concerns that iteratively retraining a generative model on its self-generated synthetic data may keep deteriorating model performance, a phenomenon often coined model collapse. In this pap…

Cited by 0SourcecodeScholar
2026

LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena

ICLR 2026poster

With the rapid progress of large language models (LLMs) trained on every available piece of data, it becomes increasingly challenging to reliably evaluate their intelligence due to potential data contamination and benchmark overfitting. To overcome these challenges, we investigate a new angle of ben…

Cited by 0SourceScholar
2026

Optimal Pricing for Data-Augmented AutoML Marketplaces

ICML 2026poster

Data markets promise to unlock data value by matching data suppliers with ML consumers. However, market design involves addressing intricate challenges, including data pricing, fairness, and robustness. We propose a pragmatic data-augmented AutoML market that seamlessly integrates with existing clou…

Cited by 0SourceScholar
2026

T-TAMER: Provably Taming Trade-offs in ML Serving

ICLR 2026poster

As machine learning models continue to grow in size and complexity, efficient serving faces increasingly broad trade-offs spanning accuracy, latency, resource usage, and other objectives. Multi-model serving further complicates these trade-offs; for example, in cascaded models, each early-exit decis…

Cited by 0SourceScholar
2025

An Instrumental Value for Data Production and its Application to Data Pricing

ICML 2025poster

We develop a framework for capturing the instrumental value of data production processes, which accounts for two key factors: (a) the context of the agent’s decision-making; (b) how much data or information the buyer already possesses. We "micro-found" our data valuation function by establishing its…

Cited by 0SourcePDFScholar
2025

Learning Personalized Ad Impact via Contextual Reinforcement Learning under Delayed Rewards

NeurIPS 2025poster

Online advertising platforms use automated auctions to connect advertisers with potential customers, requiring effective bidding strategies to maximize profits. Accurate ad impact estimation requires considering three key factors: delayed and long-term effects, cumulative ad impacts such as reinforc…

Cited by 0SourceScholar
2025

Learning from Imperfect Human Feedback: A Tale from Corruption-Robust Dueling

ICLR 2025poster

This paper studies Learning from Imperfect Human Feedback (LIHF), addressing the potential irrationality or imperfect perception when learning from comparative human feedback. Building on evidences that human's imperfection decays over time (i.e., humans learn to improve), we cast this problem as a…

Cited by 1SourcePDFScholar
2025

Mechanism Design for Large Language Models (Extended Abstract)

IJCAI 2025

We investigate auction mechanisms for AI-generated content, focusing on applications like ad creative generation. In our model, agents' preferences over stochastically generated content are encoded as large language models (LLMs). We propose an auction format that operates on a token-by-token basis,

Cited by 0SourcePDFScholar
2025

Robust Optimization with Diffusion Models for Green Security

UAI 2025

In green security, defenders must forecast adversarial behavior-such as poaching, illegal logging, and illegal fishing-to plan effective patrols. These behavior are often highly uncertain and complex. Prior work has leveraged game theory to design robust patrol strategies to handle uncertainty, but

Cited by 0SourcePDFScholar
2025

Sample Complexity of Linear Regression Models for Opinion Formation in Networks

AAAI 2025technical

Consider public health officials aiming to spread awareness about a new vaccine in a community interconnected by a social network. How can they distribute information with minimal resources, so as to avoid polarization and ensure community-wide convergence of opinion? To tackle such challenges, we i…

2024

Bandits Meet Mechanism Design to Combat Clickbait in Online Recommendation

ICLR 2024spotlight

We study a strategic variant of the multi-armed bandit problem, which we coin the strategic click-bandit. This model is motivated by applications in online recommendation where the choice of recommended items depends on both the click-through rates and the post-click rewards. Like in classical bandi…

Cited by 8SourcePDFScholar
2024

Human vs. Generative AI in Content Creation Competition: Symbiosis or Conflict?

ICML 2024poster

The advent of generative AI (GenAI) technology produces a transformative impact on the content creation landscape, offering alternative approaches to produce diverse, good-quality content across media, thereby reshaping online ecosystems but also raising concerns about market over-saturation and the…

Cited by 15SourcePDFScholar
2024

Incentivized Truthful Communication for Federated Bandits

ICLR 2024poster

To enhance the efficiency and practicality of federated bandit learning, recent advances have introduced incentives to motivate communication among clients, where a client participates only when the incentive offered by the server outweighs its participation cost. However, existing incentive mechani…

Cited by 1SourcePDFScholar
2024

Intrinsic Robustness of Prophet Inequality to Strategic Reward Signaling

NeurIPS 2024poster

Prophet inequality concerns a basic optimal stopping problem and states that simple threshold stopping policies --- i.e., accepting the first reward larger than a certain threshold --- can achieve tight $\frac{1}{2}$-approximation to the optimal prophet value. Motivated by its economic applications,…

Cited by 0SourcePDFScholar
2024

Multi-Sender Persuasion: A Computational Perspective

ICML 2024poster

We consider *multiple senders* with informational advantage signaling to convince a single self-interested actor to take certain actions. Generalizing the seminal *Bayesian Persuasion* framework, such settings are ubiquitous in computational economics, multi-agent learning, and machine learning with…

Cited by 11SourcePDFScholar
2024

Unveiling User Satisfaction and Creator Productivity Trade-Offs in Recommendation Platforms

NeurIPS 2024poster

On User-Generated Content (UGC) platforms, recommendation algorithms significantly impact creators' motivation to produce content as they compete for algorithmically allocated user traffic. This phenomenon subtly shapes the volume and diversity of the content pool, which is crucial for the platform'…

Cited by 4SourcePDFScholar
2023

Follow-ups Also Matter: Improving Contextual Bandits via Post-serving Contexts

NeurIPS 2023spotlight

Standard contextual bandit problem assumes that all the relevant contexts are observed before the algorithm chooses an arm. This modeling paradigm, while useful, often falls short when dealing with problems in which additional valuable contexts can be observed after arm selection. For example, conte…

Cited by 1SourcePDFScholar
2023

How Bad is Top-$K$ Recommendation under Competing Content Creators?

ICML 2023oral

This study explores the impact of content creators' competition on user welfare in recommendation platforms, as well as the long-term dynamics of relevance-driven recommendations. We establish a model of creator competition, under the setting where the platform uses a top-$K$ recommendation policy,…

Cited by 29SourcePDFScholar
2023

Rethinking Incentives in Recommender Systems: Are Monotone Rewards Always Beneficial?

NeurIPS 2023poster

The past decade has witnessed the flourishing of a new profession as media content creators, who rely on revenue streams from online content recommendation platforms. The reward mechanism employed by these platforms creates a competitive environment among creators which affects their production choi…

Cited by 16SourcePDFScholar
2023

The Economics of Machine Learning

IJCAI 2023poster

This survey overviews a new research agenda on the economics of machine learning, pursued at the Strategic IntelliGence for Machine Agent (SIGMA) Lab at UChicago. This overall research agenda has two themes: machine learning for economics and, conversely, economics for machine learning (ML)…

Cited by 0SourcePDFScholar
2022

CS-Shapley: Class-wise Shapley Values for Data Valuation in Classification

NeurIPS 2022accept

Data valuation, or the valuation of individual datum contributions, has seen growing interest in machine learning due to its demonstrable efficacy for tasks such as noisy label detection. In particular, due to the desirable axiomatic properties, several Shapley value approximations have been propose…

2022

First-Order Convex Fitting and Its Application to Economics and Optimization

AAAI 2022technical

This paper studies a function fitting problem which we coin first-order convex fitting (FCF): given any two vector sequences x1, ..., xT and p1, ..., pT, when is it possible to efficiently construct a convex function f(x) that ``fits'' the two sequences in the first-order sense, i.e, the (sub)g…

Cited by 3SourcePDFScholar
2022

Incrementality Bidding via Reinforcement Learning under Mixed and Delayed Rewards

NeurIPS 2022accept

Incrementality, which measures the causal effect of showing an ad to a potential customer (e.g. a user in an internet platform) versus not, is a central object for advertisers in online advertising platforms. This paper investigates the problem of how an advertiser can learn to optimize the biddin…

Cited by 2SourcePDFScholar
2022

Information design for multiple independent and self-interested defenders: Work less, pay off more

UAI 2022poster

This paper studies the problem of information design in a general security game setting in which multiple independent self-interested defenders attempt to provide protection simultaneously on the same set of important targets against an unknown attacker. A principal, who can be one of the defenders,…

Cited by 0SourcePDFScholar
2022

Inverse Game Theory for Stackelberg Games: the Blessing of Bounded Rationality

NeurIPS 2022accept

Optimizing strategic decisions (a.k.a. computing equilibrium) is key to the success of many non-cooperative multi-agent applications. However, in many real-world situations, we may face the exact opposite of this game-theoretic problem --- instead of prescribing equilibrium of a given game, we may d…

Cited by 13SourcePDFScholar
2022

Learning from a Learning User for Optimal Recommendations

ICML 2022spotlight

In real-world recommendation problems, especially those with a formidably large item space, users have to gradually learn to estimate the utility of any fresh recommendations from their experience about previously consumed items. This in turn affects their interaction dynamics with the system and ca…

Cited by 6SourcePDFScholar
2022

Learning the Optimal Recommendation from Explorative Users

AAAI 2022technical

We propose a new problem setting to study the sequential interactions between a recommender system and a user. Instead of assuming the user is omniscient, static, and explicit, as the classical practice does, we sketch a more realistic user behavior model, under which the user: 1) rejects recommenda…

Cited by 9SourcePDFScholar
2022

Saving Stochastic Bandits from Poisoning Attacks via Limited Data Verification

AAAI 2022technical

This paper studies bandit algorithms under data poisoning attacks in a bounded reward setting. We consider a strong attacker model in which the attacker can observe both the selected actions and their corresponding rewards, and can contaminate the rewards with additive noise. We show that any bandit…

Cited by 16SourcePDFScholar
2022

The Strange Role of Information Asymmetry in Auctions—Does More Accurate Value Estimation Benefit a Bidder?

AAAI 2022technical

We study the second-price auction in which bidders have asymmetric information regarding the item’s value. Each bidder’s value for the item depends on a private component and a public component. While each bidder observes their own private component, they hold different and asymmetric information ab…

Cited by 4SourcePDFScholar
2022

Understanding the Limits of Poisoning Attacks in Episodic Reinforcement Learning

IJCAI 2022poster

To understand the security threats to reinforcement learning (RL) algorithms, this paper studies poisoning attacks to manipulate any order-optimal learning algorithm towards a targeted policy in episodic RL and examines the potential damage of two natural types of poisoning attacks, i.e., the manip…

Cited by 23SourcePDFScholar
2021

(Almost) Free Incentivized Exploration from Decentralized Learning Agents

NeurIPS 2021poster

Incentivized exploration in multi-armed bandits (MAB) has witnessed increasing interests and many progresses in recent years, where a principal offers bonuses to agents to do explorations on her behalf. However, almost all existing studies are confined to temporary myopic agents. In this work, we br…

2021

Diffusion Source Identification on Networks with Statistical Confidence

ICML 2021spotlight

Diffusion source identification on networks is a problem of fundamental importance in a broad class of applications, including controlling the spreading of rumors on social media, identifying a computer virus over cyber networks, or identifying the disease center during epidemiology. Though this pro…

Cited by 12SourcePDFScholar
2020

Collapsing Bandits and Their Application to Public Health Intervention

NeurIPS 2020poster

We propose and study Collapsing Bandits, a new restless multi-armed bandit (RMAB) setting in which each arm follows a binary-state Markovian process with a special structure: when an arm is played, the state is fully observed, thus“collapsing” any uncertainty, but when an arm is passive, no observa…