← Search

Yuyang Deng

12 accepted papers

2026

Neyman-Pearson Classification under Both Null and Alternative Distributions Shift

ICLR 2026poster

We consider the problem of transfer learning in Neyman–Pearson classification, where the objective is to minimize the error w.r.t. a distribution $\mu_1$, subject to the constraint that the error w.r.t. a distribution $\mu_0$ remains below a prescribed threshold. While transfer learning has been ext…

Cited by 0SourceScholar
2025

Stochastic Compositional Minimax Optimization with Provable Convergence Guarantees

AISTATS 2025poster

Stochastic compositional minimax problems are prevalent in machine learning, yet there exist only limited established findings on the convergence of this class of problems. In this paper, we propose a formal definition of the stochastic compositional minimax problem, which involves optimizing a mini…

Cited by 0SourceScholar
2024

On the Generalization Ability of Unsupervised Pretraining

AISTATS 2024poster

Recent advances in unsupervised learning have shown that unsupervised pre-training, followed by fine-tuning, can improve model generalization. However, a rigorous understanding of how the representation function learned on an unlabeled dataset affects the generalization of the fine-tuned model is la…

Cited by 6SourcePDFScholar
2023

Distributed Personalized Empirical Risk Minimization

NeurIPS 2023poster

This paper advocates a new paradigm Personalized Empirical Risk Minimization (PERM) to facilitate learning from heterogeneous data sources without imposing stringent constraints on computational resources shared by participating devices. In PERM, we aim at learning a distinct model for each clie…

Cited by 5SourcePDFScholar
2023

Mixture Weight Estimation and Model Prediction in Multi-source Multi-target Domain Adaptation

NeurIPS 2023poster

We consider a problem of learning a model from multiple sources with the goal to perform well on a new target distribution. Such problem arises in learning with data collected from multiple sources (e.g. crowdsourcing) or learning in distributed systems, where the data can be highly heterogeneous.…

Cited by 3SourcePDFScholar
2023

Understanding Deep Gradient Leakage via Inversion Influence Functions

NeurIPS 2023poster

Deep Gradient Leakage (DGL) is a highly effective attack that recovers private training images from gradient vectors. This attack casts significant privacy challenges on distributed learning from clients with sensitive data, where clients are required to share gradients. Defending against such att…

2022

Local SGD Optimizes Overparameterized Neural Networks in Polynomial Time

AISTATS 2022poster

In this paper we prove that Local (S)GD (or FedAvg) can optimize deep neural networks with Rectified Linear Unit (ReLU) activation function in polynomial time. Despite the established convergence theory of Local SGD on optimizing general smooth functions in communication-efficient distributed optimi…

Cited by 16SourcePDFScholar
2022

Tight Analysis of Extra-gradient and Optimistic Gradient Methods For Nonconvex Minimax Problems

NeurIPS 2022accept

Despite the established convergence theory of Optimistic Gradient Descent Ascent (OGDA) and Extragradient (EG) methods for the convex-concave minimax problems, little is known about the theoretical guarantees of these methods in nonconvex settings. To bridge this gap, for the first time, this paper…

Cited by 16SourcePDFScholar
2021

Local Stochastic Gradient Descent Ascent: Convergence Analysis and Communication Efficiency

AISTATS 2021poster

Local SGD is a promising approach to overcome the communication overhead in distributed learning by reducing the synchronization frequency among worker nodes. Despite the recent theoretical advances of local SGD in empirical risk minimization, the efficiency of its counterpart in minimax optimizatio…

Cited by 78SourcePDFScholar