← Search

Aleksandar Botev

9 accepted papers

2023

Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation

ICLR 2023poster

Skip connections and normalisation layers form two standard architectural components that are ubiquitous for the training of Deep Neural Networks (DNNs), but whose precise roles are poorly understood. Recent approaches such as Deep Kernel Shaping have made progress towards reducing our reliance on t…

Cited by 37SourcePDFScholar
2022

Deep Learning without Shortcuts: Shaping the Kernel with Tailored Rectifiers

ICLR 2022poster

Training very deep neural networks is still an extremely challenging task. The common solution is to use shortcut connections and normalization layers, which are both crucial ingredients in the popular ResNet architecture. However, there is strong evidence to suggest that ResNets behave more like en…

2021

SyMetric: Measuring the Quality of Learnt Hamiltonian Dynamics Inferred from Vision

NeurIPS 2021poster

A recently proposed class of models attempts to learn latent dynamics from high-dimensional observations, like images, using priors informed by Hamiltonian mechanics. While these models have important potential applications in areas like robotics or autonomous driving, there is currently no good way…

2021

Which priors matter? Benchmarking models for learning latent dynamics

NeurIPS 2021poster

Learning dynamics is at the heart of many important applications of machine learning (ML), such as robotics and autonomous driving. In these settings, ML algorithms typically need to reason about a physical system using high dimensional observations, such as images, without access to the underlying…

Cited by 33SourcecodeScholar
2020

Hamiltonian Generative Networks

ICLR 2020spotlight

The Hamiltonian formalism plays a central role in classical and quantum physics. Hamiltonians are the main tool for modelling the continuous time evolution of systems with conserved quantities, and they come equipped with many useful properties, like time reversibility and smooth interpolation in ti…

Cited by 264SourceScholar
2018

Online Structured Laplace Approximations for Overcoming Catastrophic Forgetting

NeurIPS 2018poster

We introduce the Kronecker factored online Laplace approximation for overcoming catastrophic forgetting in neural networks. The method is grounded in a Bayesian online learning framework, where we recursively approximate the posterior after every task with a Gaussian, leading to a quadratic penalty…

2017

Complementary Sum Sampling for Likelihood Approximation in Large Scale Classification

AISTATS 2017poster

We consider training probabilistic classifiers in the case that the number of classes is too large to perform exact normalisation over all classes. We show that the source of high variance in standard sampling approximations is due to simply not including the correct class of the datapoint into the…

Cited by 34SourcePDFScholar