← Search

Masaaki Imaizumi

20 accepted papers

2026

Demystifying MaskGIT Sampler and Beyond: Adaptive Order Selection in Masked Diffusion

ICML 2026poster

Masked diffusion models have shown promising performance in generating high-quality samples in a wide range of domains, but accelerating their sampling process remains relatively underexplored. To investigate efficient samplers for masked diffusion, this paper theoretically analyzes the MaskGIT samp…

Cited by 0SourceScholar
2026

Dichotomy of Feature Learning and Unlearning: Fast-Slow Analysis on Neural Networks with Stochastic Gradient Descent

ICML 2026poster

The dynamics of gradient-based training in neural networks often exhibit nontrivial structures; hence, understanding them remains a central challenge in theoretical machine learning. In particular, a concept of *feature* **un***learning*, in which a neural network progressively loses previously lear…

Cited by 0SourceScholar
2026

Fast Escape, Slow Convergence: Learning Dynamics of Phase Retrieval under Power-Law Data

ICLR 2026oral

Scaling laws describe how learning performance improves with data, compute, or training time, and have become a central theme in modern deep learning. We study this phenomenon in a canonical nonlinear model: phase retrieval with anisotropic Gaussian inputs whose covariance spectrum follows a power l…

Cited by 0SourceScholar
2026

SONA: Learning Conditional, Unconditional, and Matching-Aware Discriminator

ICLR 2026poster

Deep generative models have made significant advances in generating complex content, yet conditional generation remains a fundamental challenge. Existing conditional generative adversarial networks often struggle to balance the dual objectives of assessing authenticity and conditional alignment of i…

Cited by 0SourcecodeScholar
2026

Spectral Gradient Descent Mitigates Anisotropy-Driven Misalignment: A Case Study in Phase Retrieval

ICML 2026poster

Spectral gradient methods, such as the Muon optimizer, modify gradient updates by preserving directional information while discarding scale, and have shown strong empirical performance in deep learning. We investigate the mechanisms underlying these gains through a dynamical analysis of a nonlinear …

Cited by 0SourceScholar
2025

Distillation of Discrete Diffusion through Dimensional Correlations

ICML 2025poster

Diffusion models have demonstrated exceptional performances in various fields of generative modeling, but suffer from slow sampling speed due to their iterative nature. While this issue is being addressed in continuous domains, discrete diffusion models face unique challenges, particularly in captur…

2025

Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor Programs

NeurIPS 2025poster

In modern theoretical analyses of neural networks, the infinite-width limit is often invoked to justify Gaussian approximations of neuron preactivations (e.g., via neural network Gaussian processes or Tensor Programs). However, these Gaussian-based asymptotic theories have so far been unable to capt…

Cited by 0SourceScholar
2025

Learning a Single Index Model from Anisotropic Data with Vanilla Stochastic Gradient Descent

AISTATS 2025poster

We investigate the problem of learning a Single Index Model (SIM)---a popular model for studying the ability of neural networks to learn features---from anisotropic Gaussian inputs by training a neuron using vanilla Stochastic Gradient Descent (SGD). While the isotropic case has been extensively stu…

Cited by 0SourceScholar
2025

Optimal Dynamic Regret by Transformers for Non-Stationary Reinforcement Learning

NeurIPS 2025poster

Transformers have demonstrated exceptional performance across a wide range of domains. While their ability to perform reinforcement learning in-context has been established both theoretically and empirically, their behavior in non-stationary environments remains less understood. In this study, we ad…

Cited by 0SourceScholar
2024

SAN: Inducing Metrizability of GAN with Discriminative Normalized Linear Layer

ICLR 2024poster

Generative adversarial networks (GANs) learn a target probability distribution by optimizing a generator and a discriminator with minimax objectives. This paper addresses the question of whether such optimization actually provides the generator with gradients that make its distribution close to the…

2023

Unified Perspective on Probability Divergence via the Density-Ratio Likelihood: Bridging KL-Divergence and Integral Probability Metrics

AISTATS 2023poster

This paper provides a unified perspective for the Kullback-Leibler (KL)-divergence and the integral probability metrics (IPMs) from the perspective of maximum likelihood density-ratio estimation (DRE). Both the KL-divergence and the IPMs are widely used in various fields in applications such as gene…

2022

Learning Causal Models from Conditional Moment Restrictions by Importance Weighting

ICLR 2022spotlight

We consider learning causal relationships under conditional moment restrictions. Unlike causal inference under unconditional moment restrictions, conditional moment restrictions pose serious challenges for causal inference. To address this issue, we propose a method that transforms conditional momen…

Cited by 9SourcePDFScholar
2021

Improved generalization bounds of group invariant / equivariant deep networks via quotient feature spaces

UAI 2021poster

Numerous invariant (or equivariant) neural networks have succeeded in handling the invariant data such as point clouds and graphs. However, a generalization theory for the neural networks has not been well developed, because several essential factors for the theory, such as network size and margin d…

Cited by 46SourcePDFScholar
2020

On Random Subsampling of Gaussian Process Regression: A Graphon-Based Analysis

AISTATS 2020poster

In this paper, we study random subsampling of Gaussian process regression, one of the simplest approximation baselines, from a theoretical perspective. Although subsampling discards a large part of training data, we show provable guarantees on the accuracy of the predictive mean/variance and its gen…

Cited by 25SourcePDFScholar
2018

Statistically Efficient Estimation for Non-Smooth Probability Densities

AISTATS 2018poster

We investigate statistical efficiency of estimators for non-smooth density functions. The density estimation problem appears in various situations, and it is intensively used in statistics and machine learning. The statistical efficiencies of estimators, i.e., their convergence rates, play a central…

Cited by 0SourcePDFScholar
2017

On Tensor Train Rank Minimization : Statistical Efficiency and Scalable Algorithm

NeurIPS 2017poster

Tensor train (TT) decomposition provides a space-efficient representation for higher-order tensors. Despite its advantage, we face two crucial limitations when we apply the TT decomposition to machine learning problems: the lack of statistical theory and of scalable algorithms. In this paper, we add…

Cited by 44SourcePDFScholar