← Search

Mihai Nica

5 accepted papers

2024

Dynamic Sparse Training with Structured Sparsity

ICLR 2024poster

Dynamic Sparse Training (DST) methods achieve state-of-the-art results in sparse neural network training, matching the generalization of dense models while enabling sparse training and inference. Although the resulting models are highly sparse and theoretically less computationally expensive, achiev…

2023

The Tilted Variational Autoencoder: Improving Out-of-Distribution Detection

ICLR 2023poster

A problem with using the Gaussian distribution as a prior for the variational autoencoder (VAE) is that the set on which Gaussians have high probability density is small as the latent dimension increases. This is an issue because VAEs try to attain both a high likelihood with respect to a prior dist…

Cited by 15SourcePDFScholar
2022

The Neural Covariance SDE: Shaped Infinite Depth-and-Width Networks at Initialization

NeurIPS 2022accept

The logit outputs of a feedforward neural network at initialization are conditionally Gaussian, given a random covariance matrix defined by the penultimate layer. In this work, we study the distribution of this random matrix. Recent work has shown that shaping the activation function as network dept…

Cited by 38SourcePDFScholar
2021

The future is log-Gaussian: ResNets and their infinite-depth-and-width limit at initialization

NeurIPS 2021poster

Theoretical results show that neural networks can be approximated by Gaussian processes in the infinite-width limit. However, for fully connected networks, it has been previously shown that for any fixed network width, $n$, the Gaussian approximation gets worse as the network depth, $d$, increases.…

Cited by 45SourcePDFScholar