← Search

Ard A. Louis

5 accepted papers

2026

Closed-form $\ell_r$ norm scaling with data for overparameterized linear regression and diagonal linear networks under $\ell_p$ bias

ICLR 2026poster

For overparameterized linear regression with isotropic Gaussian design and minimum-$\ell_p$ interpolator $p\in(1,2]$, we give a unified, high-probability characterization for the scaling of the family of parameter norms $ \\{ \lVert \widehat{w_p} \rVert_r \\}_{r \in [1,p]} $ with sample size. We s…

Cited by 0SourcecodeScholar
2026

Decoupling Dynamical Richness from Representation Learning: Towards Practical Measurement

ICLR 2026poster

Dynamic feature transformation (the rich regime) does not always align with predictive performance (better representation), yet accuracy is often used as a proxy for richness, limiting analysis of their relationship. We propose a computationally efficient, performance-independent metric of richness…

Cited by 0SourceScholar
2024

An exactly solvable model for emergence and scaling laws in the multitask sparse parity problem

NeurIPS 2024poster

Deep learning models can exhibit what appears to be a sudden ability to solve a new problem as training time, training data, or model size increases, a phenomenon known as emergence. In this paper, we present a framework where each new ability (a skill) is represented as a basis function. We solve…

Cited by 0SourcePDFScholar
2024

Double-Descent Curves in Neural Networks: A New Perspective Using Gaussian Processes

AAAI 2024technical

Double-descent curves in neural networks describe the phenomenon that the generalisation error initially descends with increasing parameters, then grows after reaching an optimal number of parameters which is less than the number of data points, but then descends again in the overparameterized regim…

Cited by 10SourcePDFScholar
2019

Deep learning generalizes because the parameter-function map is biased towards simple functions

ICLR 2019poster

Deep neural networks (DNNs) generalize remarkably well without explicit regularization even in the strongly over-parametrized regime where classical learning theory would instead predict that they would severely overfit. While many proposals for some kind of implicit regularization have been made…

Cited by 253SourcePDFScholar