← Search

Jerry Zhu

10 accepted papers

2025

A Cramér–von Mises Approach to Incentivizing Truthful Data Sharing

NeurIPS 2025poster

Modern data marketplaces and data sharing consortia increasingly rely on incentive mechanisms to encourage agents to contribute data. However, schemes that reward agents based on the quantity of submitted data are vulnerable to manipulation, as agents may submit fabricated or low-quality data to inf…

Cited by 0SourceScholar
2025

Collaborative Mean Estimation Among Heterogeneous Strategic Agents: Individual Rationality, Fairness, and Truthful Contribution

ICML 2025poster

We study a collaborative learning problem where $m$ agents aim to estimate a vector $\mu =(\mu_1,\ldots,\mu_d)\in \mathbb{R}^d$ by sampling from associated univariate normal distributions $(\mathcal{N}(\mu_k, \sigma^2))\_{k\in[d]}$. Agent $i$ incurs a cost $c_{i,k}$ to sample from $\mathcal{N}(\mu_k…

Cited by 0SourcePDFScholar
2025

Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?

NeurIPS 2025spotlight

Multi-Agent Debate (MAD) has emerged as a promising paradigm for improving the performance of large language models through collaborative reasoning. Despite recent advances, the key factors driving MAD’s effectiveness remain unclear. In this work, we disentangle MAD into two key components–Majority…

Cited by 0SourcecodeScholar
2024

Minimally Modifying a Markov Game to Achieve Any Nash Equilibrium and Value

ICML 2024poster

We study the game modification problem, where a benevolent game designer or a malevolent adversary modifies the reward function of a zero-sum Markov game so that a target deterministic or stochastic policy profile becomes the unique Markov perfect Nash equilibrium and has a value within a target ran…

2024

On the Complexity of Teaching a Family of Linear Behavior Cloning Learners

NeurIPS 2024poster

We study optimal teaching for a family of Behavior Cloning learners that learn using a linear hypothesis class. In this setup, a knowledgeable teacher can demonstrate a dataset of state and action tuples and is required to teach an optimal policy to an entire family of BC learners using the smallest…

Cited by 0SourcePDFScholar
2023

Dream the Impossible: Outlier Imagination with Diffusion Models

NeurIPS 2023poster

Utilizing auxiliary outlier datasets to regularize the machine learning model has demonstrated promise for out-of-distribution (OOD) detection and safe prediction. Due to the labor intensity in data collection and cleaning, automating outlier data generation has been a long-desired alternative. Desp…

2022

Provable Defense against Backdoor Policies in Reinforcement Learning

NeurIPS 2022accept

We propose a provable defense mechanism against backdoor policies in reinforcement learning under subspace trigger assumption. A backdoor policy is a security threat where an adversary publishes a seemingly well-behaved policy which in fact allows hidden triggers. During deployment, the adversary ca…

2021

Policy Gradient Bayesian Robust Optimization for Imitation Learning

ICML 2021spotlight

The difficulty in specifying rewards for many real-world problems has led to an increased focus on learning rewards from human feedback, such as demonstrations. However, there are often many different reward functions that explain the human feedback, leaving agents with uncertainty over what the tru…

Cited by 26SourcePDFScholar