← Search

Jorn Peters

2 accepted papers

2022

FP8 Quantization: The Power of the Exponent

NeurIPS 2022accept

When quantizing neural networks for efficient inference, low-bit integers are the go-to format for efficiency. However, low-bit floating point numbers have an extra degree of freedom, assigning some bits to work on an exponential scale instead. This paper in-depth investigates this benefit of the fl…

2019

Integer Discrete Flows and Lossless Compression

NeurIPS 2019poster

Lossless compression methods shorten the expected representation size of data without loss of information, using a statistical model. Flow-based models are attractive in this setting because they admit exact likelihood optimization, which is equivalent to minimizing the expected number of bits per m…