← Search

Yiming Ying

27 accepted papers

2026

Statistical Consistency and Generalization of Contrastive Representation Learning

ICML 2026poster

Contrastive representation learning (CRL) underpins many modern foundation models. Despite recent theoretical progress, existing analyses suffer from several key limitations: (i) the statistical consistency of CRL remains poorly understood; (ii) available generalization bounds deteriorate as the num…

Cited by 0SourceScholar
2025

How does Labeling Error Impact Contrastive Learning? A Perspective from Data Dimensionality Reduction

ICML 2025poster

In recent years, contrastive learning has achieved state-of-the-art performance in the territory of self-supervised representation learning. Many previous works have attempted to provide the theoretical understanding underlying the success of contrastive learning. Almost all of them rely on a defau…

Cited by 0SourcePDFScholar
2025

On Discriminative Probabilistic Modeling for Self-Supervised Representation Learning

ICLR 2025poster

We study the discriminative probabilistic modeling on a continuous domain for the data prediction task of (multimodal) self-supervised representation learning. To address the challenge of computing the integral in the partition function for each anchor data, we leverage the multiple importance sampl…

2025

Optimal Rates for Generalization of Gradient Descent for Deep ReLU Classification

NeurIPS 2025poster

Recent advances have significantly improved our understanding of the generalization performance of gradient descent (GD) methods in deep neural networks. A natural and fundamental question is whether GD can achieve generalization rates comparable to the minimax optimal rates established in the kerne…

Cited by 0SourceScholar
2024

Stability and Generalization of Stochastic Compositional Gradient Descent Algorithms

ICML 2024poster

Many machine learning tasks can be formulated as a stochastic compositional optimization (SCO) problem such as reinforcement learning, AUC maximization and meta-learning, where the objective function involves a nested composition associated with an expectation. Although many studies have been devote…

Cited by 2SourcePDFScholar
2023

Generalization Analysis for Contrastive Representation Learning

ICML 2023poster

Recently, contrastive learning has found impressive success in advancing the state of the art in solving various machine learning tasks. However, the existing generalization analysis is very limited or even not meaningful. In particular, the existing generalization error bounds depend linearly on th…

Cited by 11SourcePDFScholar
2023

Label Distributionally Robust Losses for Multi-class Classification: Consistency, Robustness and Adaptivity

ICML 2023poster

We study a family of loss functions named label-distributionally robust (LDR) losses for multi-class classification that are formulated from distributionally robust optimization (DRO) perspective, where the uncertainty in the given label information are modeled and captured by taking the worse case…

2023

Minimax AUC Fairness: Efficient Algorithm with Provable Convergence

AAAI 2023technical

The use of machine learning models in consequential decision making often exacerbates societal inequity, in particular yielding disparate impact on members of marginalized groups defined by race and gender. The area under the ROC curve (AUC) is widely used to evaluate the performance of a scoring fu…

2023

Three-Way Trade-Off in Multi-Objective Learning: Optimization, Generalization and Conflict-Avoidance

NeurIPS 2023poster

Multi-objective learning (MOL) often arises in emerging machine learning problems when multiple learning criteria or tasks need to be addressed. Recent works have developed various _dynamic weighting_ algorithms for MOL, including MGDA and its variants, whose central idea is to find an update direc…

2022

Differentially private SGDA for minimax problems

UAI 2022poster

Stochastic gradient descent ascent (SGDA) and its variants have been the workhorse for solving minimax problems. However, in contrast to the well-studied stochastic gradient descent (SGD) with differential privacy (DP) constraints, there is little work on understanding the generalization (utility…

2022

Stability and Generalization Analysis of Gradient Methods for Shallow Neural Networks

NeurIPS 2022accept

While significant theoretical progress has been achieved, unveiling the generalization mystery of overparameterized neural networks still remains largely elusive. In this paper, we study the generalization behavior of shallow neural networks (SNNs) by leveraging the concept of algorithmic stability…

Cited by 21SourcePDFScholar
2022

Stability and Generalization for Markov Chain Stochastic Gradient Methods

NeurIPS 2022accept

Recently there is a large amount of work devoted to the study of Markov chain stochastic gradient methods (MC-SGMs) which mainly focus on their convergence analysis for solving minimization problems. In this paper, we provide a comprehensive generalization analysis of MC-SGMs for both minimization…

Cited by 21SourcePDFScholar
2021

Distributionally Robust Optimization for Deep Kernel Multiple Instance Learning

AISTATS 2021poster

Multiple Instance Learning (MIL) provides a promising solution to many real-world problems, where labels are only available at the bag level but missing for instances due to a high labeling cost. As a powerful Bayesian non-parametric model, Gaussian Processes (GP) have been extended from classical s…

2021

Federated Deep AUC Maximization for Hetergeneous Data with a Constant Communication Complexity

ICML 2021spotlight

Deep AUC (area under the ROC curve) Maximization (DAM) has attracted much attention recently due to its great potential for imbalanced data classification. However, the research on Federated Deep AUC Maximization (FDAM) is still limited. Compared with standard federated learning (FL) approaches that…

2021

Simple Stochastic and Online Gradient Descent Algorithms for Pairwise Learning

NeurIPS 2021poster

Pairwise learning refers to learning tasks where the loss function depends on a pair of instances. It instantiates many important machine learning tasks such as bipartite ranking and metric learning. A popular approach to handle streaming data in pairwise learning is an online gradient descent (OG…

2021

Stability and Differential Privacy of Stochastic Gradient Descent for Pairwise Learning with Non-Smooth Loss

AISTATS 2021poster

Pairwise learning has recently received increasing attention since it subsumes many important machine learning tasks (e.g. AUC maximization and metric learning) into a unifying framework. In this paper, we give the first-ever-known stability and generalization analysis of stochastic gradient descent…

Cited by 23SourcePDFScholar
2021

Stability and Generalization of Stochastic Gradient Methods for Minimax Problems

ICML 2021oral

Many machine learning problems can be formulated as minimax problems such as Generative Adversarial Networks (GANs), AUC maximization and robust estimation, to mention but a few. A substantial amount of studies are devoted to studying the convergence behavior of their stochastic gradient-type algori…

2019

Stochastic Iterative Hard Thresholding for Graph-structured Sparsity Optimization

ICML 2019oral

Stochastic optimization algorithms update models with cheap per-iteration costs sequentially, which makes them amenable for large-scale data analysis. Such algorithms have been widely studied for structured sparse models where the sparsity information is very specific, e.g., convex sparsity-inducing…

2016

Fast Convergence of Online Pairwise Learning Algorithms

AISTATS 2016poster

Pairwise learning usually refers to a learning task which involves a loss function depending on pairs of examples, among which most notable ones are bipartite ranking, metric learning and AUC maximization. In this paper, we focus on online learning algorithms for pairwise learning problems without…

Cited by 20SourcePDFScholar