← Search

Gabriel Pereyra

2 accepted papers

2018

Large scale distributed neural network training through online distillation

ICLR 2018poster

Techniques such as ensembling and distillation promise model quality improvements when paired with almost any base model. However, due to increased test-time cost (for ensembles) and increased complexity of the training pipeline (for distillation), these techniques are challenging to use in industri…

Cited by 535SourcePDFScholar
2016

Batch normalized recurrent neural networks

ICASSP 2016accepted

Recurrent Neural Networks (RNNs) are powerful models for sequential data that have the potential to learn long-term dependencies. However, they are computationally expensive to train and difficult to parallelize. Recent work has shown that normalizing intermediate representations of neural networks…

Cited by 0SourceScholar