← Search

Berfin Simsek

9 accepted papers

2025

Flat Channels to Infinity in Neural Loss Landscapes

NeurIPS 2025poster

The loss landscapes of neural networks contain minima and saddle points that may be connected in flat regions or appear in isolation. We identify and characterize a special structure in the loss landscape: channels along which the loss decreases extremely slowly, while the output weights of at least…

Cited by 0SourceScholar
2025

Learning Gaussian Multi-Index Models with Gradient Flow: Time Complexity and Directional Convergence

AISTATS 2025poster

This work focuses on the gradient flow dynamics of a neural network model that uses correlation loss to approximate a multi-index function on high-dimensional standard Gaussian data. Specifically, the multi-index function we consider is a sum of neurons $f^*(x) = \sum_{j=1}^k \sigma^*(v_j^T x)$ whe…

Cited by 0SourceScholar
2025

Loss Landscape of Shallow ReLU-like Neural Networks: Stationary Points, Saddle Escape, and Network Embedding

ICLR 2025poster

In this paper, we study the loss landscape of one-hidden-layer neural networks with ReLU-like activation functions trained with the empirical squared loss using gradient descent (GD). We identify the stationary points of such networks, which significantly slow down loss decrease during training. To…

Cited by 0SourcePDFScholar
2024

Expand-and-Cluster: Parameter Recovery of Neural Networks

ICML 2024poster

Can we identify the weights of a neural network by probing its input-output mapping? At first glance, this problem seems to have many solutions because of permutation, overparameterisation and activation function symmetries. Yet, we show that the incoming weight vector of each neuron is identifiable…

2023

Should Under-parameterized Student Networks Copy or Average Teacher Weights?

NeurIPS 2023poster

Any continuous function $f^*$ can be approximated arbitrarily well by a neural network with sufficiently many neurons $k$. We consider the case when $f^*$ itself is a neural network with one hidden layer and $k$ neurons. Approximating $f^*$ with a neural network with $n< k$ neurons can thus be seen…

2021

Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances

ICML 2021spotlight

We study how permutation symmetries in overparameterized multi-layer neural networks generate ‘symmetry-induced’ critical points. Assuming a network with $ L $ layers of minimal widths $ r_1^*, \ldots, r_{L-1}^* $ reaches a zero-loss minimum at $ r_1^*! \cdots r_{L-1}^*! $ isolated points that are p…

2020

Implicit Regularization of Random Feature Models

ICML 2020poster

Random Features (RF) models are used as efficient parametric approximations of kernel methods. We investigate, by means of random matrix theory, the connection between Gaussian RF models and Kernel Ridge Regression (KRR). For a Gaussian RF model with $P$ features, $N$ data points, and a ridge $\lamb…

Cited by 108SourcePDFScholar
2020

Kernel Alignment Risk Estimator: Risk Prediction from Training Data

NeurIPS 2020poster

We study the risk (i.e. generalization error) of Kernel Ridge Regression (KRR) for a kernel $K$ with ridge $\lambda>0$ and i.i.d. observations. For this, we introduce two objects: the Signal Capture Threshold (SCT) and the Kernel Alignment Risk Estimator (KARE). The SCT $\vartheta_{K,\lambda}$ is a…

Cited by 72SourcePDFScholar