← Search

Shun-ichi Amari

5 accepted papers

2021

When does preconditioning help or hurt generalization?

ICLR 2021poster

While second order optimizers such as natural gradient descent (NGD) often speed up optimization, their effect on generalization has been called into question. This work presents a more nuanced view on how the \textit{implicit bias} of optimizers affects the comparison of generalization properties.…

Cited by 50SourcePDFScholar
2019

Fisher Information and Natural Gradient Learning in Random Deep Networks

AISTATS 2019poster

The parameter space of a deep neural network is a Riemannian manifold, where the metric is defined by the Fisher information matrix. The natural gradient method uses the steepest descent direction in a Riemannian manifold, but it requires inversion of the Fisher matrix, however, which is practicall…

Cited by 49SourcePDFScholar
2019

Interpolating between Optimal Transport and MMD using Sinkhorn Divergences

AISTATS 2019poster

Comparing probability distributions is a fundamental problem in data sciences. Simple norms and divergences such as the total variation and the relative entropy only compare densities in a point-wise manner and fail to capture the geometric nature of the problem. In sharp contrast, Maximum Mean Disc…

2019

The Normalization Method for Alleviating Pathological Sharpness in Wide Neural Networks

NeurIPS 2019poster

Normalization methods play an important role in enhancing the performance of deep learning while their theoretical understandings have been limited. To theoretically elucidate the effectiveness of normalization, we quantify the geometry of the parameter space determined by the Fisher information mat…

Cited by 52SourcePDFScholar
2019

Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach

AISTATS 2019poster

The Fisher information matrix (FIM) is a fundamental quantity to represent the characteristics of a stochastic model, including deep neural networks (DNNs). The present study reveals novel statistics of FIM that are universal among a wide class of DNNs. To this end, we use random weights and large w…

Cited by 157SourcePDFScholar