← Search

Ryo Karakida

13 accepted papers

2025

Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor Programs

NeurIPS 2025poster

In modern theoretical analyses of neural networks, the infinite-width limit is often invoked to justify Gaussian approximations of neuron preactivations (e.g., via neural network Gaussian processes or Tensor Programs). However, these Gaussian-based asymptotic theories have so far been unable to capt…

Cited by 0SourceScholar
2025

Local Loss Optimization in the Infinite Width: Stable Parameterization of Predictive Coding Networks and Target Propagation

ICLR 2025poster

Local learning, which trains a network through layer-wise local targets and losses, has been studied as an alternative to backpropagation (BP) in neural computation. However, its algorithms often become more complex or require additional hyperparameters due to the locality, making it challenging to…

Cited by 0SourcePDFScholar
2024

On the Parameterization of Second-Order Optimization Effective towards the Infinite Width

ICLR 2024poster

Second-order optimization has been developed to accelerate the training of deep neural networks and it is being applied to increasingly larger-scale models. In this study, towards training on further larger scales, we identify a specific parameterization for second-order optimization that promotes f…

Cited by 5SourcePDFScholar
2023

Understanding Gradient Regularization in Deep Learning: Efficient Finite-Difference Computation and Implicit Bias

ICML 2023poster

Gradient regularization (GR) is a method that penalizes the gradient norm of the training loss during training. While some studies have reported that GR can improve generalization performance, little attention has been paid to it from the algorithmic perspective, that is, the algorithms of GR that e…

Cited by 13SourcePDFScholar
2022

Learning curves for continual learning in neural networks: Self-knowledge transfer and forgetting

ICLR 2022poster

Sequential training from task to task is becoming one of the major objects in deep learning applications such as continual learning and transfer learning. Nevertheless, it remains unclear under what conditions the trained model's performance improves or deteriorates. To deepen our understanding of s…

Cited by 23SourcePDFScholar
2021

The Spectrum of Fisher Information of Deep Networks Achieving Dynamical Isometry

AISTATS 2021poster

The Fisher information matrix (FIM) is fundamental to understanding the trainability of deep neural nets (DNN), since it describes the parameter space’s local metric. We investigate the spectral distribution of the conditional FIM, which is the FIM given a single sample, by focusing on fully-connect…

Cited by 7SourcePDFScholar
2020

Understanding Approximate Fisher Information for Fast Convergence of Natural Gradient Descent in Wide Neural Networks

NeurIPS 2020oral

Natural Gradient Descent (NGD) helps to accelerate the convergence of gradient descent dynamics, but it requires approximations in large-scale deep neural networks because of its high computational cost. Empirical studies have confirmed that some NGD methods with approximate Fisher information conve…

2019

Fisher Information and Natural Gradient Learning in Random Deep Networks

AISTATS 2019poster

The parameter space of a deep neural network is a Riemannian manifold, where the metric is defined by the Fisher information matrix. The natural gradient method uses the steepest descent direction in a Riemannian manifold, but it requires inversion of the Fisher matrix, however, which is practicall…

Cited by 49SourcePDFScholar
2019

The Normalization Method for Alleviating Pathological Sharpness in Wide Neural Networks

NeurIPS 2019poster

Normalization methods play an important role in enhancing the performance of deep learning while their theoretical understandings have been limited. To theoretically elucidate the effectiveness of normalization, we quantify the geometry of the parameter space determined by the Fisher information mat…

Cited by 52SourcePDFScholar
2019

Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach

AISTATS 2019poster

The Fisher information matrix (FIM) is a fundamental quantity to represent the characteristics of a stochastic model, including deep neural networks (DNNs). The present study reveals novel statistics of FIM that are universal among a wide class of DNNs. To this end, we use random weights and large w…

Cited by 157SourcePDFScholar