← Search

Mohamed Osama Ahmed

7 accepted papers

2024

Memory Efficient Neural Processes via Constant Memory Attention Block

ICML 2024poster

Neural Processes (NPs) are popular meta-learning methods for efficiently modelling predictive uncertainty. Recent state-of-the-art methods, however, leverage expensive attention mechanisms, limiting their applications, particularly in low-resource settings. In this work, we propose Constant Memory A…

2024

Tree Cross Attention

ICLR 2024poster

Cross Attention is a popular method for retrieving information from a set of context tokens for making predictions. At inference time, for each prediction, Cross Attention scans the full set of $\mathcal{O}(N)$ tokens. In practice, however, often only a small subset of tokens are required for good p…

2023

Latent Bottlenecked Attentive Neural Processes

ICLR 2023poster

Neural Processes (NPs) are popular methods in meta-learning that can estimate predictive uncertainty on target datapoints by conditioning on a context dataset. Previous state-of-the-art method Transformer Neural Processes (TNPs) achieve strong performance but require quadratic computation with respe…

2023

Towards Better Selective Classification

ICLR 2023poster

We tackle the problem of Selective Classification where the objective is to achieve the best performance on a predetermined ratio (coverage) of the dataset. Recent state-of-the-art selective methods come with architectural changes either via introducing a separate selection head or an extra abstenti…

2022

Monotonicity regularization: Improved penalties and novel applications to disentangled representation learning and robust classification

UAI 2022poster

We study settings where gradient penalties are used alongside risk minimization with the goal of obtaining predictors satisfying different notions of monotonicity. Specifically, we present two sets of contributions. In the first part of the paper, we show that different choices of penalties define t…

2015

StopWasting My Gradients: Practical SVRG

NeurIPS 2015poster

We present and analyze several strategies for improving the performance ofstochastic variance-reduced gradient (SVRG) methods. We first show that theconvergence rate of these methods can be preserved under a decreasing sequenceof errors in the control variate, and use this to derive variants of SVRG…

Cited by 169SourcePDFScholar