← Search

Dayal Singh Kalra

4 accepted papers

2025

(How) Can Transformers Predict Pseudo-Random Numbers?

ICML 2025poster

Transformers excel at discovering patterns in sequential data, yet their fundamental limitations and learning mechanisms remain crucial topics of investigation. In this paper, we study the ability of Transformers to learn pseudo-random number sequences from linear congruential generators (LCGs), def…

Cited by 0SourcePDFScholar
2025

Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos

ICLR 2025poster

In gradient descent dynamics of neural networks, the top eigenvalue of the Hessian of the loss (sharpness) displays a variety of robust phenomena throughout training. This includes early time regimes where the sharpness may decrease during early periods of training (sharpness reduction), and later t…

Cited by 7SourcePDFScholar
2023

Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and width

NeurIPS 2023poster

We systematically analyze optimization dynamics in deep neural networks (DNNs) trained with stochastic gradient descent (SGD) and study the effect of learning rate $\eta$, depth $d$, and width $w$ of the neural network. By analyzing the maximum eigenvalue $\lambda^H_t$ of the Hessian of the loss, wh…

Cited by 13SourcePDFScholar