ICLR 2023poster7 citations

Sequential Learning of Neural Networks for Prequential MDL

Jorg Bornschein, Yazhe Li, Marcus Hutter

Abstract

Minimum Description Length (MDL) provides a framework and an objective for principled model evaluation. It formalizes Occam's Razor and can be applied to data from non-stationary sources. In the prequential formulation of MDL, the objective is to minimize the cumulative next-step log-loss when sequentially going through the data and using previous observations for parameter estimation. It thus closely resembles a continual- or online-learning problem. In this study, we evaluate approaches for computing prequential description lengths for image classification datasets with neural networks. Considering the computational cost, we find that online-learning with rehearsal has favorable performance compared to the previously widely used block-wise estimation. We propose forward-calibration to better align the models predictions with the empirical observations and introduce replay-streams, a minibatch incremental training technique to efficiently implement approximate random replay while avoiding large in-memory replay buffers. As a result, we present description lengths for a suite of image classification datasets that improve upon previously reported results by large margins.

mdlcontinual-learningdeep-learning
BibTeX
@inproceedings{
bornschein2023sequential,
title={Sequential Learning of Neural Networks for Prequential {MDL}},
author={Jorg Bornschein and Yazhe Li and Marcus Hutter},
booktitle={The Eleventh International Conference on Learning Representations },
year={2023},
url={https://openreview.net/forum?id=dMMPUvNSYJr}
}
Sequential Learning of Neural Networks for Prequential MDL · ICLR 2023