← Search

Moritz Haas

4 accepted papers

2025

On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling

NeurIPS 2025spotlight

Scaling limits, such as infinite-width limits, serve as promising theoretical tools to study large-scale models. However, it is widely believed that existing infinite-width theory does not faithfully explain the behavior of practical networks, especially those trained in *standard parameterization*…

Cited by 0SourceScholar
2024

$\boldsymbol{\mu}\mathbf{P^2}$: Effective Sharpness Aware Minimization Requires Layerwise Perturbation Scaling

NeurIPS 2024poster

Sharpness Aware Minimization (SAM) enhances performance across various neural architectures and datasets. As models are continually scaled up to improve performance, a rigorous understanding of SAM’s scaling behaviour is paramount. To this end, we study the infinite-width limit of neural networks tr…

Cited by 0SourcePDFScholar
2024

On Feature Learning in Structured State Space Models

NeurIPS 2024poster

This paper studies the scaling behavior of state-space models (SSMs) and their structured variants, such as Mamba, that have recently arisen in popularity as alternatives to transformer-based neural network architectures. Specifically, we focus on the capability of SSMs to learn features as their ne…

Cited by 2SourcePDFScholar
2023

Mind the spikes: Benign overfitting of kernels and neural networks in fixed dimension

NeurIPS 2023poster

The success of over-parameterized neural networks trained to near-zero training error has caused great interest in the phenomenon of benign overfitting, where estimators are statistically consistent even though they interpolate noisy training data. While benign overfitting in fixed dimension has bee…