← Search

Guido Montufar

26 accepted papers

2025

Demystifying Topological Message-Passing with Relational Structures: A Case Study on Oversquashing in Simplicial Message-Passing

ICLR 2025poster

Topological deep learning (TDL) has emerged as a powerful tool for modeling higher-order interactions in relational data. However, phenomena such as oversquashing in topological message-passing remain understudied and lack theoretical analysis. We propose a unifying axiomatic framework that bridges…

Cited by 0SourcePDFScholar
2025

Implicit Bias of Mirror Flow for Shallow Neural Networks in Univariate Regression

ICLR 2025spotlight

We examine the implicit bias of mirror flow in least squares error regression with wide and shallow neural networks. For a broad class of potential functions, we show that mirror flow exhibits lazy training and has the same implicit bias as ordinary gradient flow when the network width tends to infi…

Cited by 0SourcePDFScholar
2025

Zero-Shot Context Generalization in Reinforcement Learning from Few Training Contexts

NeurIPS 2025poster

Deep reinforcement learning (DRL) has achieved remarkable success across multiple domains, including competitive games, natural language processing, and robotics. Despite these advancements, policies trained via DRL often struggle to generalize to evaluation environments with different parameters. T…

Cited by 0SourceScholar
2024

Benign overfitting in leaky ReLU networks with moderate input dimension

NeurIPS 2024spotlight

The problem of benign overfitting asks whether it is possible for a model to perfectly fit noisy training data and still generalize well. We study benign overfitting in two-layer leaky ReLU networks trained with the hinge loss on a binary classification task. We consider input data which can be deco…

Cited by 4SourcePDFScholar
2024

Bounds for the smallest eigenvalue of the NTK for arbitrary spherical data of arbitrary dimension

NeurIPS 2024poster

Bounds on the smallest eigenvalue of the neural tangent kernel (NTK) are a key ingredient in the analysis of neural network optimization and memorization. However, existing results require distributional assumptions on the data and are limited to a high-dimensional setting, where the input dimension…

Cited by 3SourcePDFScholar
2023

Characterizing the spectrum of the NTK via a power series expansion

ICLR 2023poster

Under mild conditions on the network initialization we derive a power series expansion for the Neural Tangent Kernel (NTK) of arbitrarily deep feedforward networks in the infinite width limit. We provide expressions for the coefficients of this power series which depend on both the Hermite coefficie…

2023

Critical Points and Convergence Analysis of Generative Deep Linear Networks Trained with Bures-Wasserstein Loss

ICML 2023poster

We consider a deep matrix factorization model of covariance matrices trained with the Bures-Wasserstein distance. While recent works have made advances in the study of the optimization problem for overparametrized low-rank matrix approximation, much emphasis has been placed on discriminative setting…

2023

Expected Gradients of Maxout Networks and Consequences to Parameter Initialization

ICML 2023poster

We study the gradients of a maxout network with respect to inputs and parameters and obtain bounds for the moments depending on the architecture and the parameter distribution. We observe that the distribution of the input-output Jacobian depends on the input, which complicates a stable parameter in…

2023

FoSR: First-order spectral rewiring for addressing oversquashing in GNNs

ICLR 2023poster

Graph neural networks (GNNs) are able to leverage the structure of graph data by passing messages along the edges of the graph. While this allows GNNs to learn features depending on the graph structure, for certain graph topologies it leads to inefficient information propagation and a problem known…

2022

Learning Curves for Gaussian Process Regression with Power-Law Priors and Targets

ICLR 2022poster

We characterize the power-law asymptotics of learning curves for Gaussian process regression (GPR) under the assumption that the eigenspectrum of the prior and the eigenexpansion coefficients of the target function follow a power law. Under similar assumptions, we leverage the equivalence between GP…

Cited by 17SourcePDFScholar
2022

Spectral Bias Outside the Training Set for Deep Networks in the Kernel Regime

NeurIPS 2022accept

We provide quantitative bounds measuring the $L^2$ difference in function space between the trajectory of a finite-width network trained on finitely many samples from the idealized kernel dynamics of infinite width and infinite data. An implication of the bounds is that the network is biased to lea…

2022

The Geometry of Memoryless Stochastic Policy Optimization in Infinite-Horizon POMDPs

ICLR 2022poster

We consider the problem of finding the best memoryless stochastic policy for an infinite-horizon partially observable Markov decision process (POMDP) with finite state and action spaces with respect to either the discounted or mean reward criterion. We show that the (discounted) state-action frequen…

2021

How Framelets Enhance Graph Neural Networks

ICML 2021spotlight

This paper presents a new approach for assembling graph neural networks based on framelet transforms. The latter provides a multi-scale representation for graph-structured data. We decompose an input graph into low-pass and high-pass frequencies coefficients for network training, which then defines…

2021

Weisfeiler and Lehman Go Cellular: CW Networks

NeurIPS 2021poster

Graph Neural Networks (GNNs) are limited in their expressive power, struggle with long-range interactions and lack a principled way to model higher-order structures. These problems can be attributed to the strong coupling between the computational graph and the input graph structure. The recently pr…

2020

Optimization Theory for ReLU Neural Networks Trained with Normalization Layers

ICML 2020poster

The current paradigm of deep neural networks has been successful in part due to the use of normalization layers. Normalization layers like Batch Normalization, Layer Normalization and Weight Normalization are ubiquitous in practice as they improve the generalization performance and training speed of…

Cited by 36SourcePDFScholar