← Search

Parthe Pandit

13 accepted papers

2025

Emergence in non-neural models: grokking modular arithmetic via average gradient outer product

ICML 2025oral

Neural networks trained to solve modular arithmetic tasks exhibit grokking, a phenomenon where the test accuracy starts improving long after the model achieves 100% training accuracy in the training process. It is often taken as an example of "emergence", where model ability manifests sharply throug…

Cited by 6SourcePDFScholar
2025

Fast Training of Large Kernel Models with Delayed Projections

NeurIPS 2025spotlight

Classical kernel machines have historically faced significant challenges in scaling to large datasets and model sizes—a key ingredient that has driven the success of neural networks. In this paper, we present a new methodology for building kernel machines that can scale efficiently with both data si…

Cited by 0SourcecodeScholar
2024

On the Nyström Approximation for Preconditioning in Kernel Machines

AISTATS 2024poster

Kernel methods are a popular class of nonlinear predictive models in machine learning. Scalable algorithms for learning kernel models need to be iterative in nature, but convergence can be slow due to poor conditioning. Spectral preconditioning is an important tool to speed-up the convergence of suc…

Cited by 4SourcePDFScholar
2022

Benign, Tempered, or Catastrophic: Toward a Refined Taxonomy of Overfitting

NeurIPS 2022accept

The practical success of overparameterized neural networks has motivated the recent scientific study of \emph{interpolating methods}-- learning methods which are able fit their training data perfectly. Empirically, certain interpolating methods can fit noisy training data without catastrophically ba…

Cited by 45SourcePDFScholar
2022

Instability and Local Minima in GAN Training with Kernel Discriminators

NeurIPS 2022accept

Generative Adversarial Networks (GANs) are a widely-used tool for generative modeling of complex data. Despite their empirical success, the training of GANs is not fully understood due to the joint training of the generator and discriminator. This paper analyzes these joint dynamics when the true s…

Cited by 12SourcePDFScholar
2021

Implicit Bias of Linear RNNs

ICML 2021spotlight

Contemporary wisdom based on empirical studies suggests that standard recurrent neural networks (RNNs) do not perform well on tasks requiring long-term memory. However, RNNs’ poor ability to capture long-term dependencies has not been fully understood. This paper provides a rigorous explanation of t…

Cited by 13SourcePDFScholar
2020

Generalization Error of Generalized Linear Models in High Dimensions

ICML 2020poster

At the heart of machine learning lies the question of generalizability of learned rules over previously unseen data. While over-parameterized models based on neural networks are now ubiquitous in machine learning applications, our understanding of their generalization capabilities is incomplete and…

Cited by 64SourcePDFScholar
2020

Matrix Inference and Estimation in Multi-Layer Models

NeurIPS 2020poster

We consider the problem of estimating the input and hidden variables of a stochastic multi-layer neural network from an observation of the output. The hidden variables in each layer are represented as matrices with statistical interactions along both rows as well as columns. This problem applies to…

2019

Sparse Multivariate Bernoulli Processes in High Dimensions

AISTATS 2019poster

We consider the problem of estimating the parameters of a multivariate Bernoulli process with auto-regressive feedback in the high-dimensional setting where the number of samples available is much less than the number of parameters. This problem arises in learning interconnections of networks of dyn…

Cited by 6SourcePDFScholar
2018

Plug-in Estimation in High-Dimensional Linear Inverse Problems: A Rigorous Analysis

NeurIPS 2018poster

Estimating a vector $\mathbf{x}$ from noisy linear measurements $\mathbf{Ax+w}$ often requires use of prior knowledge or structural constraints on $\mathbf{x}$ for accurate reconstruction. Several recent works have considered combining linear least-squares estimation with a generic or plug-in ``deno…

Cited by 75SourcePDFScholar
2015

Structural segmentation of Hindustani concert audio with posterior features

ICASSP 2015accepted

Structural segmentation of music involves identifying boundaries between homogenous regions where the homogeneity involves one or more musical dimensions, and therefore depends on the musical genre. In this work, we address the segmentation of Hindustani instrumental concert recordings at the highes…

Cited by 0SourceScholar