← Search

Xueru Zhang

24 accepted papers

2026

Achieving Fairness Without Harm via Selective Demographic Experts

AAAI 2026technical

As machine learning systems become increasingly integrated into human-centered domains such as healthcare, ensuring fairness while maintaining high predictive performance is critical. Existing bias mitigation techniques often impose a trade-off between fairness and accuracy, inadvertently degrading

Cited by 2SourcePDFScholar
2026

Stabilizing Self-Consuming Diffusion Models with Latent Space Filtering

AAAI 2026technical

As synthetic data proliferates across the Internet, it is often reused to train successive generations of generative models. This creates a "self-consuming loop" that can lead to training instability or *model collapse*. Common strategies to address the issue---such as accumulating historical traini

Cited by 0SourcePDFScholar
2026

When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming Loop

ICML 2026poster

Foundation models are increasingly trained on synthetic data generated by prior model iterations rather than exclusively on real data. This *self-consuming* training paradigm can lead to model collapse, divergence, or bias amplification. Recent work (Ferbach et al., 2024) shows that incorporating hu…

Cited by 0SourceScholar
2025

DroughtSet: Understanding Drought Through Spatial-Temporal Learning

AAAI 2025technical

Drought is one of the most destructive and expensive natural disasters, severely impacting natural resources and risks by depleting water resources and diminishing agricultural yields. Under climate change, accurately predicting drought is critical for mitigating drought-induced risks. However, the…

2025

Open-Set Heterogeneous Domain Adaptation: Theoretical Analysis and Algorithm

AAAI 2025technical

Domain adaptation (DA) tackles the issue of distribution shift by learning a model from a source domain that generalizes to a target domain. However, most existing DA methods are designed for scenarios where the source and target domain data lie within the same feature space, which limits their appl…

Cited by 0SourcePDFScholar
2025

The Boundaries of Fair AI in Medical Image Prognosis: A Causal Perspective

NeurIPS 2025poster

As machine learning (ML) algorithms are increasingly used in medical image analysis, concerns have emerged about their potential biases against certain social groups. Although many approaches have been proposed to ensure the fairness of ML models, most existing works focus only on medical image diag…

Cited by 0SourceScholar
2024

Automating Data Annotation under Strategic Human Agents: Risks and Potential Solutions

NeurIPS 2024poster

As machine learning (ML) models are increasingly used in social domains to make consequential decisions about humans, they often have the power to reshape data distributions. Humans, as strategic agents, continuously adapt their behaviors in response to the learning system. As populations change dyn…

2024

Performative Federated Learning: A Solution to Model-Dependent and Heterogeneous Distribution Shifts

AAAI 2024technical

We consider a federated learning (FL) system consisting of multiple clients and a server, where the clients aim to collaboratively learn a common decision model from their distributed data. Unlike the conventional FL framework that assumes the client's data is static, we consider scenarios where the…

2024

Privacy-Aware Randomized Quantization via Linear Programming

UAI 2024poster

Differential privacy mechanisms such as the Gaussian or Laplace mechanism have been widely used in data analytics for preserving individual privacy. However, they are mostly designed for continuous outputs and are unsuitable for scenarios where discrete values are necessary. Although various quantiz…

2022

Fairness Interventions as (Dis)Incentives for Strategic Manipulation

ICML 2022spotlight

Although machine learning (ML) algorithms are widely used to make decisions about individuals in various domains, concerns have arisen that (1) these algorithms are vulnerable to strategic manipulation and "gaming the algorithm"; and (2) ML decisions may exhibit bias against certain social groups. E…

Cited by 26SourcePDFScholar
2021

Fair Sequential Selection Using Supervised Learning Models

NeurIPS 2021poster

We consider a selection problem where sequentially arrived applicants apply for a limited number of positions/jobs. At each time step, a decision maker accepts or rejects the given applicant using a pre-trained supervised learning model until all the vacant positions are filled. In this paper, we di…

2021

Improving Fairness and Privacy in Selection Problems

AAAI 2021technical

Supervised learning models have been increasingly used for making decisions about individuals in applications such as hiring, lending, and college admission. These models may inherit pre-existing biases from training datasets and discriminate against protected attributes (e.g., race or gender). In a…

Cited by 37SourcePDFScholar
2020

How do fair decisions fare in long-term qualification?

NeurIPS 2020poster

Although many fairness criteria have been proposed for decision making, their long-term impact on the well-being of a population remains unclear. In this work, we study the dynamics of population qualification and algorithmic decisions under a partially observed Markov decision problem setting. By c…

2019

Group Retention when Using Machine Learning in Sequential Decision Making: the Interplay between User Dynamics and Fairness

NeurIPS 2019poster

Machine Learning (ML) models trained on data from multiple demographic groups can inherit representation disparity (Hashimoto et al., 2018) that may exist in the data: the model may be less favorable to groups contributing less to the training process; this in turn can degrade population retention i…

Cited by 68SourcePDFScholar
2018

Improving the Privacy and Accuracy of ADMM-Based Distributed Algorithms

ICML 2018oral

Alternating direction method of multiplier (ADMM) is a popular method used to design distributed versions of a machine learning algorithm, whereby local computations are performed on local data with the output exchanged among neighbors in an iterative fashion. During this iterative process the leaka…

Cited by 119SourcePDFScholar