← Search

Hossein Taheri

8 accepted papers

2026

On the Theory of Continual Learning with Gradient Descent for Neural Networks

ICML 2026poster

Continual learning, the ability of a model to adapt to an ongoing sequence of tasks without forgetting earlier ones, is a central goal of artificial intelligence. To better understand its underlying mechanisms, we study the limitations of continual learning in a tractable yet representative setting.…

Cited by 0SourceScholar
2025

Sharper Guarantees for Learning Neural Network Classifiers with Gradient Methods

ICLR 2025poster

In this paper, we study the data-dependent convergence and generalization behavior of gradient methods for neural networks with smooth activation. Our first result is a novel bound on the excess risk of deep networks trained by the logistic loss via an alogirthmic stability analysis. Compared to pre…

Cited by 0SourcePDFScholar
2023

Fast Convergence in Learning Two-Layer Neural Networks with Separable Data

AAAI 2023technical

Normalized gradient descent has shown substantial success in speeding up the convergence of exponentially-tailed loss functions (which includes exponential and logistic losses) on linear classifiers with separable data. In this paper, we go beyond linear models by studying normalized GD on two-lay…

Cited by 2SourcePDFScholar
2021

Fundamental Limits of Ridge-Regularized Empirical Risk Minimization in High Dimensions

AISTATS 2021poster

Despite the popularity of Empirical Risk Minimization (ERM) algorithms, a theory that explains their statistical properties in modern high-dimensional regimes is only recently emerging. We characterize for the first time the fundamental limits on the statistical accuracy of convex ridge-regularized…

Cited by 48SourcePDFScholar
2020

Quantized Decentralized Stochastic Learning over Directed Graphs

ICML 2020poster

We consider a decentralized stochastic learning problem where data points are distributed among computing nodes communicating over a directed graph. As the model size gets large, decentralized learning faces a major bottleneck that is the heavy communication load due to each node transmitting large…

Cited by 69SourcePDFScholar
2020

Sharp Asymptotics and Optimal Performance for Inference in Binary Models

AISTATS 2020poster

We study convex empirical risk minimization for high-dimensional inference in binary models. Our first result sharply predicts the statistical performance of such estimators in the linear asymptotic regime under isotropic Gaussian features. Importantly, the predictions hold for a wide class of conve…

Cited by 41SourcePDFScholar
2019

Robust and Communication-Efficient Collaborative Learning

NeurIPS 2019poster

We consider a decentralized learning problem, where a set of computing nodes aim at solving a non-convex optimization problem collaboratively. It is well-known that decentralized optimization schemes face two major system bottlenecks: stragglers' delay and communication overhead. In this paper, we t…

Cited by 126SourcePDFScholar