← Search

Benoit Dherin

7 accepted papers

2026

Equivalence of Context and Parameter Updates in Modern Transformer Blocks

ICML 2026oral

Recent research has established that the impact of context in a vanilla transformer can be represented implicitly by forming a token-dependent, rank-1 patch to its MLP weights. This work extends that foundational theory to the diverse architectures of modern Large Language Models. We first demonstra…

Cited by 0SourceScholar
2024

Deep Fusion: Efficient Network Training via Pre-trained Initializations

ICML 2024poster

Training deep neural networks for large language models (LLMs) remains computationally very expensive. To mitigate this, network growing algorithms offer potential cost savings, but their underlying mechanisms are poorly understood. In this paper, we propose a theoretical framework using backward er…

Cited by 4SourcePDFScholar
2024

The Impact of Geometric Complexity on Neural Collapse in Transfer Learning

NeurIPS 2024poster

Many of the recent advances in computer vision and language models can be attributed to the success of transfer learning via the pre-training of large foundation models. However, a theoretical framework which explains this empirical success is incomplete and remains an active area of research. Flatn…

Cited by 0SourcePDFScholar
2022

Why neural networks find simple solutions: The many regularizers of geometric complexity

NeurIPS 2022accept

In many contexts, simpler models are preferable to more complex models and the control of this model complexity is the goal for many methods in machine learning such as regularization, hyperparameter tuning and architecture design. In deep learning, it has been difficult to understand the underlying…

Cited by 39SourcePDFScholar
2021

On the Origin of Implicit Regularization in Stochastic Gradient Descent

ICLR 2021poster

For infinitesimal learning rates, stochastic gradient descent (SGD) follows the path of gradient flow on the full batch loss function. However moderately large learning rates can achieve higher test accuracies, and this generalization benefit is not explained by convergence bounds, since the learnin…

Cited by 248SourcePDFScholar