← Search

Pedro Henrique Pamplona Savarese

5 accepted papers

2023

Accelerated Training via Incrementally Growing Neural Networks using Variance Transfer and Learning Rate Adaptation

NeurIPS 2023poster

We develop an approach to efficiently grow neural networks, within which parameterization and optimization strategies are designed by considering their effects on the training dynamics. Unlike existing growing methods, which follow simple replication heuristics or utilize auxiliary gradient-based l…

Cited by 7SourcePDFScholar
2022

Not All Bits have Equal Value: Heterogeneous Precisions via Trainable Noise

NeurIPS 2022accept

We study the problem of training deep networks while quantizing parameters and activations into low-precision numeric representations, a setting central to reducing energy consumption and inference time of deployed models. We propose a method that learns different precisions, as measured by bits in…

Cited by 8SourcePDFScholar
2021

Growing Efficient Deep Networks by Structured Continuous Sparsification

ICLR 2021oral

We develop an approach to growing deep network architectures over the course of training, driven by a principled combination of accuracy and sparsity objectives. Unlike existing pruning or architecture search techniques that operate on full-sized models or supernet architectures, our method can sta…

Cited by 69SourcePDFScholar
2021

Online Meta-Learning via Learning with Layer-Distributed Memory

NeurIPS 2021poster

We demonstrate that efficient meta-learning can be achieved via end-to-end training of deep neural networks with memory distributed across layers. The persistent state of this memory assumes the entire burden of guiding task adaptation. Moreover, its distributed nature is instrumental in orchestra…

Cited by 5SourcePDFScholar
2019

Convergence of Gradient Descent on Separable Data

AISTATS 2019poster

We provide a detailed study on the implicit bias of gradient descent when optimizing loss functions with strictly monotone tails, such as the logistic loss, over separable datasets. We look at two basic questions: (a) what are the conditions on the tail of the loss function under which gradient desc…

Cited by 186SourcePDFScholar