← Search

Atsushi Nitanda

36 accepted papers

2026

Provable Sample Efficiency of Curriculum Post-Training for Transformer Reasoning

ICML 2026poster

Recent curriculum techniques in the post-training stage of LLMs have been empirically observed to outperform non-curriculum approaches in improving reasoning performance, yet a principled understanding of their effectiveness and limitations remains incomplete. To bridge this gap, we develop an abstr…

Cited by 0SourceScholar
2025

Direct Distributional Optimization for Provable Alignment of Diffusion Models

ICLR 2025poster

We introduce a novel alignment method for diffusion models from distribution optimization perspectives while providing rigorous convergence guarantees. We first formulate the problem as a generic regularized loss minimization over probability distributions and directly optimize the distribution usin…

Cited by 0SourcePDFScholar
2025

Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model Ensemble

ICML 2025poster

Mean-field Langevin dynamics (MFLD) is an optimization method derived by taking the mean-field limit of noisy gradient descent for two-layer neural networks in the mean-field regime. Recently, the propagation of chaos (PoC) for MFLD has gained attention as it provides a quantitative characterization…

Cited by 0SourcePDFScholar
2025

Provable In-Context Vector Arithmetic via Retrieving Task Concepts

ICML 2025poster

In-context learning (ICL) has garnered significant attention for its ability to grasp functions/tasks from demonstrations. Recent studies suggest the presence of a latent **task/function vector** in LLMs during ICL. Merullo et al. (2024) showed that LLMs leverage this vector alongside the residual s…

Cited by 0SourcePDFScholar
2025

Statistical Analysis of the Sinkhorn Iterations for Two-Sample Schr\"{o}dinger Bridge Estimation

NeurIPS 2025poster

The Schrödinger bridge problem seeks the optimal stochastic process that connects two given probability distributions with minimal energy modification. While the Sinkhorn algorithm is widely used to solve the static optimal transport problem, a recent work (Pooladian and Niles-Weed, 2024) proposed…

Cited by 0SourceScholar
2024

Improved statistical and computational complexity of the mean-field Langevin dynamics under structured data

ICLR 2024poster

Recent works have shown that neural networks optimized by gradient-based methods can adapt to sparse or low-dimensional target functions through feature learning; an often studied target is the sparse parity function on the unit hypercube. However, such isotropic data setting does not capture the an…

Cited by 6SourcePDFScholar
2024

Koopman-based generalization bound: New aspect for full-rank weights

ICLR 2024poster

We propose a new bound for generalization of neural networks using Koopman operators. Whereas most of existing works focus on low-rank weight matrices, we focus on full-rank weight matrices. Our bound is tighter than existing norm-based bounds when the condition numbers of weight matrices are small.…

Cited by 3SourcePDFScholar
2024

Provably Transformers Harness Multi-Concept Word Semantics for Efficient In-Context Learning

NeurIPS 2024poster

Transformer-based large language models (LLMs) have displayed remarkable creative prowess and emergence capabilities. Existing empirical studies have revealed a strong connection between these LLMs' impressive emergence abilities and their in-context learning (ICL) capacity, allowing them to solve n…

Cited by 0SourcePDFScholar
2024

Why is parameter averaging beneficial in SGD? An objective smoothing perspective

AISTATS 2024poster

It is often observed that stochastic gradient descent (SGD) and its variants implicitly select a solution with good generalization performance; such implicit bias is often characterized in terms of the sharpness of the minima. Kleinberg et al. (2018) connected this bias with the smoothing effect of…

Cited by 0SourcePDFScholar
2023

Convergence of mean-field Langevin dynamics: time-space discretization, stochastic gradient, and variance reduction

NeurIPS 2023spotlight

The mean-field Langevin dynamics (MFLD) is a nonlinear generalization of the Langevin dynamics that incorporates a distribution-dependent drift, and it naturally arises from the optimization of two-layer neural networks via (noisy) gradient descent. Recent works have shown that MFLD globally minimiz…

Cited by 16SourcePDFScholar
2023

Feature learning via mean-field Langevin dynamics: classifying sparse parities and beyond

NeurIPS 2023poster

Neural network in the mean-field regime is known to be capable of \textit{feature learning}, unlike the kernel (NTK) counterpart. Recent works have shown that mean-field neural networks can be globally optimized by a noisy gradient descent update termed the \textit{mean-field Langevin dynamics} (MFL…

Cited by 17SourcePDFScholar
2023

Primal and Dual Analysis of Entropic Fictitious Play for Finite-sum Problems

ICML 2023poster

The entropic fictitious play (EFP) is a recently proposed algorithm that minimizes the sum of a convex functional and entropy in the space of measures --- such an objective naturally arises in the optimization of a two-layer neural network in the mean-field regime. In this work, we provide a concise…

Cited by 5SourcePDFScholar
2023

Tight and fast generalization error bound of graph embedding in metric space

ICML 2023poster

Recent studies have experimentally shown that we can achieve in non-Euclidean metric space effective and efficient graph embedding, which aims to obtain the vertices' representations reflecting the graph's structure in the metric space. Specifically, graph embedding in hyperbolic space has experimen…

Cited by 1SourcePDFScholar
2023

Uniform-in-time propagation of chaos for the mean-field gradient Langevin dynamics

ICLR 2023poster

The mean-field Langevin dynamics is characterized by a stochastic differential equation that arises from (noisy) gradient descent on an infinite-width two-layer neural network, which can be viewed as an interacting particle system. In this work, we establish a quantitative weak propagation of chaos…

Cited by 21SourcePDFScholar
2022

Particle Stochastic Dual Coordinate Ascent: Exponential convergent algorithm for mean field neural network optimization

ICLR 2022poster

We introduce Particle-SDCA, a gradient-based optimization algorithm for two-layer neural networks in the mean field regime that achieves exponential convergence rate in regularized empirical risk minimization. The proposed algorithm can be regarded as an infinite dimensional extension of Stochastic…

Cited by 15SourcePDFScholar
2022

Two-layer neural network on infinite dimensional data: global optimization guarantee in the mean-field regime

NeurIPS 2022accept

Analysis of neural network optimization in the mean-field regime is important as the setting allows for feature learning. Existing theory has been developed mainly for neural networks in finite dimensions, i.e., each neuron has a finite-dimensional parameter. However, the setting of infinite-dimensi…

Cited by 6SourcePDFScholar
2021

Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic Besov space

NeurIPS 2021spotlight

Deep learning has exhibited superior performance for various tasks, especially for high-dimensional datasets, such as images. To understand this property, we investigate the approximation and estimation ability of deep learning on {\it anisotropic Besov spaces}. The anisotropic Besov space is chara…

Cited by 85SourcePDFScholar
2021

Exponential Convergence Rates of Classification Errors on Learning with SGD and Random Features

AISTATS 2021poster

Although kernel methods are widely used in many learning problems, they have poor scalability to large datasets. To address this problem, sketching and stochastic gradient methods are the most commonly used techniques to derive computationally efficient learning algorithms. We consider solving a bin…

Cited by 3SourcePDFScholar
2021

Generalization Bounds for Graph Embedding Using Negative Sampling: Linear vs Hyperbolic

NeurIPS 2021poster

Graph embedding, which represents real-world entities in a mathematical space, has enabled numerous applications such as analyzing natural languages, social networks, biochemical networks, and knowledge bases. It has been experimentally shown that graph embedding in hyperbolic space can represent hi…

Cited by 12SourcePDFScholar
2021

Generalization Error Bound for Hyperbolic Ordinal Embedding

ICML 2021spotlight

Hyperbolic ordinal embedding (HOE) represents entities as points in hyperbolic space so that they agree as well as possible with given constraints in the form of entity $i$ is more similar to entity $j$ than to entity $k$. It has been experimentally shown that HOE can obtain representations of hiera…

Cited by 14SourcePDFScholar
2021

Optimal Rates for Averaged Stochastic Gradient Descent under Neural Tangent Kernel Regime

ICLR 2021oral

We analyze the convergence of the averaged stochastic gradient descent for overparameterized two-layer neural networks for regression problems. It was recently found that a neural tangent kernel (NTK) plays an important role in showing the global convergence of gradient-based methods under the NTK r…

Cited by 59SourcePDFScholar
2021

Particle Dual Averaging: Optimization of Mean Field Neural Network with Global Convergence Rate Analysis

NeurIPS 2021poster

We propose the particle dual averaging (PDA) method, which generalizes the dual averaging method in convex optimization to the optimization over probability distributions with quantitative runtime guarantee. The algorithm consists of an inner loop and outer loop: the inner loop utilizes the Langevin…

Cited by 22SourcePDFScholar
2021

When does preconditioning help or hurt generalization?

ICLR 2021poster

While second order optimizers such as natural gradient descent (NGD) often speed up optimization, their effect on generalization has been called into question. This work presents a more nuanced view on how the \textit{implicit bias} of optimizers affects the comparison of generalization properties.…

Cited by 50SourcePDFScholar
2020

Functional Gradient Boosting for Learning Residual-like Networks with Statistical Guarantees

AISTATS 2020poster

Recently, several studies have proposed progressive or sequential layer-wise training methods based on the boosting theory for deep neural networks. However, most studies lack the global convergence guarantees or require weak learning conditions that can be verified a posteriori after running method…

Cited by 10SourcePDFScholar
2019

Stochastic Gradient Descent with Exponential Convergence Rates of Expected Classification Errors

AISTATS 2019poster

We consider stochastic gradient descent and its averaging variant for binary classification problems in a reproducing kernel Hilbert space. In traditional analysis using a consistency property of loss functions, it is known that the expected classification error converges more slowly than the expect…

Cited by 12SourcePDFScholar
2018

Gradient Layer: Enhancing the Convergence of Adversarial Training for Generative Models

AISTATS 2018poster

We propose a new technique that boosts the convergence of training generative adversarial networks. Generally, the rate of training deep models reduces severely after multiple iterations. A key reason for this phenomenon is that a deep network is expressed using a highly non-convex finite-dimensiona…

Cited by 0SourcePDFScholar
2017

Stochastic Difference of Convex Algorithm and its Application to Training Deep Boltzmann Machines

AISTATS 2017poster

Difference of convex functions (DC) programming is an important approach to nonconvex optimization problems because these structures can be encountered in several fields. Effective optimization methods, called DC algorithms, have been developed in deterministic optimization literature. In machine le…

Cited by 37SourcePDFScholar