← Search

Tom Bäckström

7 accepted papers

2026

DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick

ICLR 2026poster

Vector quantization is common in deep models, yet its hard assignments block gradients and hinder end-to-end training. We propose DiVeQ, which treats quantization as adding an error vector that mimics the quantization distortion, keeping the forward pass hard while letting gradients flow. We also pr…

Cited by 0SourcecodeScholar
2023

Stochastic Optimization of Vector Quantization Methods in Application to Speech and Image Processing

ICASSP 2023accepted

Vector quantization (VQ) methods have been used in a wide range of applications for speech, image, and video data. While classic VQ methods often use expectation maximization, in this paper, we investigate the use of stochastic optimization employing our recently proposed noise substitution in vecto…

Cited by 0SourceScholar
2018

GMM-Based Iterative Entropy Coding for Spectral Envelopes of Speech and Audio

ICASSP 2018accepted

Spectral envelope modelling is a central part of speech and audio codecs and is traditionally based on either vector quantization or scalar quantization followed by entropy coding. To bridge the coding performance of vector quantization with the low complexity of the scalar case, we propose an itera…

Cited by 0SourceScholar
2015

Arithmetic coding of speech and audio spectra using tcx based on linear predictive spectral envelopes

ICASSP 2015accepted

Unified speech and audio codecs often use a frequency domain coding technique of the transform coded excitation (TCX) type. It is based on modeling the speech source with a linear predictor, spectral weighting by a perceptual model and entropy coding of the frequency components. While previous appro…

Cited by 0SourceScholar
2015

Finding line spectral frequencies using the fast fourier transform

ICASSP 2015accepted

Main-stream speech codecs are based on modelling the speech source by a linear predictor. An efficient domain for quantization and coding of this linear predictor is the line spectral frequency representation, where the predictor is encoded into an ordered set of frequencies that correspond to the r…

Cited by 0SourceScholar
2015

Intelligibility evaluation of speech coding standards in severe background noise and packet loss conditions

ICASSP 2015accepted

Speech intelligibility is an important aspect of speech transmission but often when speech coding standards are compared only the quality is evaluated using perceptual tests. In this study, the performance of three wideband speech coding standards, adaptive multi-rate wideband (AMR-WB), G.718, and e…

Cited by 0SourceScholar