← Search

Debarghya Ghoshdastidar

19 accepted papers

2026

Interpretable Self-Supervised Learning via Representer Landmarks and Nyström Approximation

ICML 2026poster

Self-supervised learning (SSL) effectively learns representations from massive unlabeled data, yet the resulting models typically operate as black boxes, necessitating domain-specific post-hoc explanations. We introduce KREPES, a unified framework that learns inherently interpretable representations…

Cited by 0SourceScholar
2025

Exact Certification of (Graph) Neural Networks Against Label Poisoning

ICLR 2025spotlight

Machine learning models are highly vulnerable to label flipping, i.e., the adversarial modification (poisoning) of training labels to compromise performance. Thus, deriving robustness certificates is important to guarantee that test predictions remain unaffected and to understand worst-case robustne…

2025

Infinite Width Limits of Self Supervised Neural Networks

AISTATS 2025poster

The NTK is a widely used tool in the theoretical analysis of deep learning, allowing us to look at supervised deep neural networks through the lenses of kernel regression. Recently, several works have investigated kernel models for self-supervised learning, hypothesizing that these also shed light o…

Cited by 0SourceScholar
2025

Non-Singularity of the Gradient Descent Map for Neural Networks with Piecewise Analytic Activations

NeurIPS 2025poster

The theory of training deep networks has become a central question of modern machine learning and has inspired many practical advancements. In particular, the gradient descent (GD) optimization algorithm has been extensively studied in recent years. A key assumption about GD has appeared in several…

Cited by 0SourceScholar
2025

When Can We Approximate Wide Contrastive Models with Neural Tangent Kernels and Principal Component Analysis?

AAAI 2025technical

Contrastive learning is a paradigm for learning representations from unlabelled data and several recent works have claimed that such models effectively learn spectral embeddings and show relations between (wide) contrastive models and kernel principal component analysis (PCA). However, it is not kno…

Cited by 1SourcePDFScholar
2024

Explaining Kernel Clustering via Decision Trees

ICLR 2024poster

Despite the growing popularity of explainable and interpretable machine learning, there is still surprisingly limited work on inherently interpretable clustering methods. Recently, there has been a surge of interest in explaining the classic k-means algorithm, leading to efficient algorithms that ap…

Cited by 2SourcePDFScholar
2024

Non-parametric Representation Learning with Kernels

AAAI 2024technical

Unsupervised and self-supervised representation learning has become popular in recent years for learning useful features from unlabelled data. Representation learning has been mostly developed in the neural network literature, and other models for representation learning are surprisingly unexplored.…

Cited by 7SourcePDFScholar
2023

Improved Representation Learning Through Tensorized Autoencoders

AISTATS 2023poster

The central question in representation learning is what constitutes a good or meaningful representation. In this work we argue that if we consider data with inherent cluster structures, where clusters can be characterized through different means and covariances, those data structures should be repre…

Cited by 2SourcePDFScholar
2022

Causal forecasting: generalization bounds for autoregressive models

UAI 2022poster

Despite the increasing relevance of forecasting methods, causal implications of these algorithms remain largely unexplored. This is concerning considering that, even under simplifying assumptions such as causal sufficiency, the statistical risk of a model can differ significantly from its causal ris…

2022

Graphon based Clustering and Testing of Networks: Algorithms and Theory

ICLR 2022poster

Network-valued data are encountered in a wide range of applications, and pose challenges in learning due to their complex structure and absence of vertex correspondence. Typical examples of such problems include classification or grouping of protein structures and social networks. Various methods, r…

2022

Interpolation and Regularization for Causal Learning

NeurIPS 2022accept

Recent work shows that in complex model classes, interpolators can achieve statistical generalization and even be optimal for statistical learning. However, despite increasing interest in learning models with good causal properties, there is no understanding of whether such interpolators can also ac…

Cited by 3SourcePDFScholar
2021

Learning Theory Can (Sometimes) Explain Generalisation in Graph Neural Networks

NeurIPS 2021poster

In recent years, several results in the supervised learning setting suggested that classical statistical learning-theoretic measures, such as VC dimension, do not adequately explain the performance of deep learning models which prompted a slew of work in the infinite-width and iteration regimes. How…

Cited by 66SourcePDFScholar
2021

Recovery Guarantees for Kernel-based Clustering under Non-parametric Mixture Models

AISTATS 2021poster

Despite the ubiquity of kernel-based clustering, surprisingly few statistical guarantees exist beyond settings that consider strong structural assumptions on the data generation process. In this work, we take a step towards bridging this gap by studying the statistical performance of kernel-based cl…

Cited by 3SourcePDFScholar
2019

Foundations of Comparison-Based Hierarchical Clustering

NeurIPS 2019poster

We address the classical problem of hierarchical clustering, but in a framework where one does not have access to a representation of the objects or their pairwise similarities. Instead, we assume that only a set of comparisons between objects is available, that is, statements of the form ``objects…

2015

A Provable Generalized Tensor Spectral Method for Uniform Hypergraph Partitioning

ICML 2015poster

Matrix spectral methods play an important role in statistics and machine learning, and most often the word ‘matrix’ is dropped as, by default, one assumes that similarities or affinities are measured between two points, thereby resulting in similarity matrices. However, recent challenges in computer…

Cited by 65SourcePDFScholar