← Search

Robert Ormandi

1 accepted papers

2018

Large scale distributed neural network training through online distillation

ICLR 2018poster

Techniques such as ensembling and distillation promise model quality improvements when paired with almost any base model. However, due to increased test-time cost (for ensembles) and increased complexity of the training pipeline (for distillation), these techniques are challenging to use in industri…

Cited by 535SourcePDFScholar