← Search

Blake Bordelon

27 accepted papers

2026

CompleteP for RL: Maintaining Feature Learning When Scaling Deep Reinforcement Learning

ICML 2026poster

The maximal update parameterization ($\mu P$) has been influential in supervised and unsupervised learning conditions, with fixed data distributions, owing to its ability to maintain feature learning across larger parameter scales. This parameterization facilitates more consistent learning dynamics …

Cited by 0SourceScholar
2026

Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time

ICLR 2026poster

We study in-context learning (ICL) of linear regression in a deep linear self-attention model, characterizing how performance depends on various computational and statistical resources (width, depth, number of training steps, batch size and data per context). In a joint limit where data dimension, c…

Cited by 0SourceScholar
2025

A Model of Place Field Reorganization During Reward Maximization

ICML 2025poster

When rodents learn to navigate in a novel environment, a high density of place fields emerges at reward locations, fields elongate against the trajectory, and individual fields change spatial selectivity while demonstrating stable behavior. Why place fields demonstrate these characteristic phenomena…

Cited by 1SourcePDFScholar
2025

Adaptive kernel predictors from feature-learning infinite limits of neural networks

ICML 2025poster

Previous influential work showed that infinite width limits of neural networks in the lazy training regime are described by kernel machines. Here, we show that neural networks trained in the rich infinite-width regime in two different settings are also described by kernel machines, but with data-dep…

Cited by 2SourcePDFScholar
2025

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer

ICML 2025poster

We theoretically characterize gradient descent dynamics in deep linear networks trained at large width from random initialization and on large quantities of random data. Our theory captures the ``wider is better" effect of mean-field/maximum-update parameterized networks as well as hyperparameter tr…

Cited by 1SourcePDFScholar
2025

Do Mice Grok? Glimpses of Hidden Progress in Sensory Cortex

ICLR 2025poster

Does learning of task-relevant representations stop when behavior stops changing? Motivated by recent work in machine learning and the intuitive observation that human experts continue to learn after mastery, we hypothesize that task-specific representation learning in cortex can continue, even when…

Cited by 0SourcePDFScholar
2025

Don't be lazy: CompleteP enables compute-efficient deep transformers

NeurIPS 2025poster

We study compute efficiency of LLM training when using different parameterizations, i.e., rules for adjusting model and optimizer hyperparameters (HPs) as model size changes. Some parameterizations fail to transfer optimal base HPs (such as learning rate) across changes in model depth, requiring pra…

Cited by 0SourcecodeScholar
2025

Scaling Laws for Precision

ICLR 2025oral

Low precision training and inference affect both the quality and cost of language models, but current scaling laws do not account for this. In this work, we devise "precision-aware" scaling laws for both training and inference. We propose that training in lower precision reduces the model's "effecti…

Cited by 24SourcePDFScholar
2024

Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit

ICLR 2024poster

The cost of hyperparameter tuning in deep learning has been rising with model sizes, prompting practitioners to find new tuning methods using a proxy of smaller networks. One such proposal uses $\mu$P parameterized networks, where the optimal hyperparameters for small width networks *transfer* to ne…

Cited by 28SourcePDFScholar
2024

Grokking as the transition from lazy to rich training dynamics

ICLR 2024poster

We propose that the grokking phenomenon, where the train loss of a neural network decreases much earlier than its test loss, can arise due to a neural network transitioning from lazy training dynamics to a rich, feature learning regime. To illustrate this mechanism, we study the simple setting of va…

Cited by 0SourcePDFScholar
2023

Dynamics of Finite Width Kernel and Prediction Fluctuations in Mean Field Neural Networks

NeurIPS 2023spotlight

We analyze the dynamics of finite width effects in wide but finite feature learning neural networks. Starting from a dynamical mean field theory description of infinite width deep neural network kernel and prediction dynamics, we provide a characterization of the $\mathcal{O}(1/\sqrt{\text{width}})…

2023

Feature-Learning Networks Are Consistent Across Widths At Realistic Scales

NeurIPS 2023poster

We study the effect of width on the dynamics of feature-learning neural networks across a variety of architectures and datasets. Early in training, wide neural networks trained on online data have not only identical loss curves but also agree in their point-wise test predictions throughout training.…

Cited by 31SourcePDFScholar
2023

Loss Dynamics of Temporal Difference Reinforcement Learning

NeurIPS 2023poster

Reinforcement learning has been successful across several applications in which agents have to learn to act in environments with sparse feedback. However, despite this empirical success there is still a lack of theoretical understanding of how the parameters of reinforcement learning models and the…

2023

The Influence of Learning Rule on Representation Dynamics in Wide Neural Networks

ICLR 2023top-25%

It is unclear how changing the learning rule of a deep neural network alters its learning dynamics and representations. To gain insight into the relationship between learned features, function approximation, and the learning rule, we analyze infinite-width deep networks trained with gradient descent…

Cited by 30SourcePDFScholar
2023

The Onset of Variance-Limited Behavior for Networks in the Lazy and Rich Regimes

ICLR 2023poster

For small training set sizes $P$, the generalization error of wide neural networks is well-approximated by the error of an infinite width neural network (NN), either in the kernel or mean-field/feature-learning regime. However, after a critical sample size $P^*$, we empirically find the finite-width…

2022

Capacity of Group-invariant Linear Readouts from Equivariant Representations: How Many Objects can be Linearly Classified Under All Possible Views?

ICLR 2022poster

Equivariance has emerged as a desirable property of representations of objects subject to identity-preserving transformations that constitute a group, such as translations and rotations. However, the expressivity of a representation constrained by group equivariance is still not fully understood. We…

2022

Neural Networks as Kernel Learners: The Silent Alignment Effect

ICLR 2022poster

Neural networks in the lazy training regime converge to kernel machines. Can neural networks in the rich feature learning regime learn a kernel machine with a data-dependent kernel? We demonstrate that this can indeed happen due to a phenomenon we term silent alignment, which requires that the tange…

Cited by 107SourcePDFScholar
2022

Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural Networks

NeurIPS 2022accept

We analyze feature learning in infinite-width neural networks trained with gradient flow through a self-consistent dynamical field theory. We construct a collection of deterministic dynamical order parameters which are inner-product kernels for hidden unit activations and gradients in each layer at…

Cited by 95SourcePDFScholar
2021

Efficient online inference for nonparametric mixture models

UAI 2021poster

Natural data are often well-described as belonging to latent clusters. When the number of clusters is unknown, Bayesian nonparametric (BNP) models can provide a flexible and powerful technique to model the data. However, algorithms for inference in nonparametric mixture models fail to meet two criti…

2021

Out-of-Distribution Generalization in Kernel Regression

NeurIPS 2021poster

In real word applications, data generating process for training a machine learning model often differs from what the model encounters in the test stage. Understanding how and whether machine learning models generalize under such distributional shifts have been a theoretical challenge. Here, we stud…

2020

Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural Networks

ICML 2020poster

We derive analytical expressions for the generalization performance of kernel regression as a function of the number of training samples using theoretical methods from Gaussian processes and statistical physics. Our expressions apply to wide neural networks due to an equivalence between training the…