← Search

Yingxue Zhou

6 accepted papers

2025

Loss Gradient Gaussian Width based Generalization and Optimization Guarantees

AISTATS 2025oral

Generalization and optimization guarantees on the population loss often rely on uniform convergence based analysis, typically based on the Rademacher complexity of the predictors. The rich representation power of modern models has led to concerns about this approach. In this paper, we present genera…

Cited by 0SourceScholar
2024

RecMind: Large Language Model Powered Agent For Recommendation

NAACL 2024findings

While the recommendation system (RS) has advanced significantly through deep learning, current RS approaches usually train and fine-tune models on task-specific datasets, limiting their generalizability to new recommendation tasks and their ability to leverage external knowledge due to model scale a…

Cited by 144SourcePDFScholar
2022

Stability Based Generalization Bounds for Exponential Family Langevin Dynamics

ICML 2022spotlight

Recent years have seen advances in generalization bounds for noisy stochastic algorithms, especially stochastic gradient Langevin dynamics (SGLD) based on stability (Mou et al., 2018; Li et al., 2020) and information theoretic approaches (Xu & Raginsky, 2017; Negrea et al., 2019; Steinke & Zakynthin…

Cited by 13SourcePDFScholar
2021

Bypassing the Ambient Dimension: Private SGD with Gradient Subspace Identification

ICLR 2021poster

Differentially private SGD (DP-SGD) is one of the most popular methods for solving differentially private empirical risk minimization (ERM). Due to its noisy perturbation on each gradient update, the error rate of DP-SGD scales with the ambient dimension $p$, the number of parameters in the model. S…

Cited by 128SourcePDFScholar
2020

Towards Better Generalization of Adaptive Gradient Methods

NeurIPS 2020poster

Adaptive gradient methods such as AdaGrad, RMSprop and Adam have been optimizers of choice for deep learning due to their fast training speed. However, it was recently observed that their generalization performance is often worse than that of SGD for over-parameterized neural networks. While new alg…

Cited by 25SourcePDFScholar