← Search

Xiaotong Yuan

17 accepted papers

2026

Learning to Refine: Spectral-Decoupled Iterative Refinement Framework for Precipitation Nowcasting

ICML 2026poster

Accurate precipitation nowcasting is vital for disaster mitigation, but deep learning methods suffer a key trade-off: regression models produce over-smoothed, spectrally decaying predictions that blur convective details and violate turbulence power laws; diffusion models generate realistic yet unanc…

Cited by 0SourceScholar
2026

LoPhyDA: Low-Rank Tensor and Physics Gradient Guided Diffusion for Atmospheric Data Assimilation

ICML 2026poster

Data Assimilation (DA) aims to integrate observations with model forecasts to estimate the state of dynamical systems. Despite the widespread application of diffusion-based assimilation methods, they remain constrained by the high dimensionality of atmospheric states and the reliance on imperfect st…

Cited by 0SourceScholar
2025

Optimization over Sparse Support-Preserving Sets: Two-Step Projection with Global Optimality Guarantees

ICML 2025poster

In sparse optimization, enforcing hard constraints using the $\ell_0$ pseudo-norm offers advantages like controlled sparsity compared to convex relaxations. However, many real-world applications demand not only sparsity constraints but also some extra constraints. While prior algorithms have been d…

2023

$L_2$-Uniform Stability of Randomized Learning Algorithms: Sharper Generalization Bounds and Confidence Boosting

NeurIPS 2023poster

Exponential generalization bounds with near-optimal rates have recently been established for uniformly stable algorithms~\citep{feldman2019high,bousquet2020sharper}. We seek to extend these best known high probability bounds from deterministic learning algorithms to the regime of randomized learning…

Cited by 2SourcePDFScholar
2022

On Convergence of FedProx: Local Dissimilarity Invariant Bounds, Non-smoothness and Beyond

NeurIPS 2022accept

The \FedProx~algorithm is a simple yet powerful distributed proximal point optimization method widely used for federated learning (FL) over heterogeneous data. Despite its popularity and remarkable success witnessed in practice, the theoretical understanding of FedProx is largely underinvestigated:…

Cited by 76SourcePDFScholar
2022

Zeroth-Order Hard-Thresholding: Gradient Error vs. Expansivity

NeurIPS 2022accept

$\ell_0$ constrained optimization is prevalent in machine learning, particularly for high-dimensional problems, because it is a fundamental approach to achieve sparse learning. Hard-thresholding gradient descent is a dominant technique to solve this problem. However, first-order gradients of the obj…

Cited by 7SourcePDFScholar
2021

A Theory-Driven Self-Labeling Refinement Method for Contrastive Representation Learning

NeurIPS 2021spotlight

For an image query, unsupervised contrastive learning labels crops of the same image as positives, and other image crops as negatives. Although intuitive, such a native label assignment strategy cannot reveal the underlying semantic similarity between a query and its positives and negatives,…

Cited by 13SourcePDFScholar
2021

Towards Understanding Why Lookahead Generalizes Better Than SGD and Beyond

NeurIPS 2021poster

To train networks, lookahead algorithm~\cite{zhang2019lookahead} updates its fast weights $k$ times via an inner-loop optimizer before updating its slow weights once by using the latest fast weights. Any optimizer, e.g. SGD, can serve as the inner-loop optimizer, and the derived lookahead gen…

2019

Efficient Meta Learning via Minibatch Proximal Update

NeurIPS 2019spotlight

We address the problem of meta-learning which learns a prior over hypothesis from a sample of meta-training tasks for fast adaptation on meta-testing tasks. A particularly simple yet successful paradigm for this research is model-agnostic meta-learning (MAML). Implementation and analysis of MAML, ho…

Cited by 117SourcePDFScholar
2018

New Insight into Hybrid Stochastic Gradient Descent: Beyond With-Replacement Sampling and Convexity

NeurIPS 2018poster

As an incremental-gradient algorithm, the hybrid stochastic gradient descent (HSGD) enjoys merits of both stochastic and full gradient methods for finite-sum minimization problem. However, the existing rate-of-convergence analysis for HSGD is made under with-replacement sampling (WRS) and is restr…

Cited by 29SourcePDFScholar
2016

Learning Additive Exponential Family Graphical Models via $\ell_{2,1}$-norm Regularized M-Estimation

NeurIPS 2016poster

We investigate a subclass of exponential family graphical models of which the sufficient statistics are defined by arbitrary additive forms. We propose two $\ell_{2,1}$-norm regularized maximum likelihood estimators to learn the model parameters from i.i.d. samples. The first one is a joint MLE esti…

Cited by 9SourcePDFScholar