← Search

Mufan Bill Li

3 accepted papers

2024

Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit

ICLR 2024poster

The cost of hyperparameter tuning in deep learning has been rising with model sizes, prompting practitioners to find new tuning methods using a proxy of smaller networks. One such proposal uses $\mu$P parameterized networks, where the optimal hyperparameters for small width networks *transfer* to ne…

Cited by 28SourcePDFScholar
2023

The Shaped Transformer: Attention Models in the Infinite Depth-and-Width Limit

NeurIPS 2023poster

In deep learning theory, the covariance matrix of the representations serves as a proxy to examine the network’s trainability. Motivated by the success of Transform- ers, we study the covariance matrix of a modified Softmax-based attention model with skip connections in the proportional limit of inf…

Cited by 40SourcePDFScholar
2022

The Neural Covariance SDE: Shaped Infinite Depth-and-Width Networks at Initialization

NeurIPS 2022accept

The logit outputs of a feedforward neural network at initialization are conditionally Gaussian, given a random covariance matrix defined by the penultimate layer. In this work, we study the distribution of this random matrix. Recent work has shown that shaping the activation function as network dept…

Cited by 38SourcePDFScholar