← Search

Michael Chinen

6 accepted papers

2023

LMCodec: A Low Bitrate Speech Codec with Causal Transformer Models

ICASSP 2023accepted

We introduce LMCodec, a causal neural speech codec that provides high quality audio at very low bitrates. The backbone of the system is a causal convolutional codec that encodes audio into a hierarchy of coarse-to-fine tokens using residual vector quantization. LMCodec trains a Transformer language…

Cited by 0SourceScholar
2021

Generative Speech Coding with Predictive Variance Regularization

ICASSP 2021accepted

The recent emergence of machine-learning based generative models for speech suggests a significant reduction in bit rate for speech codecs is possible. However, the performance of generative models deteriorates significantly with the distortions present in real-world input signals. We argue that thi…

Cited by 0SourceScholar
2021

Warp-Q: Quality Prediction for Generative Neural Speech Codecs

ICASSP 2021accepted

Good speech quality has been achieved using waveform matching and parametric reconstruction coders. Recently developed very low bit rate generative codecs can reconstruct high quality wideband speech with bit streams less than 3 kb/s. These codecs use a DNN with parametric input to synthesise high q…

Cited by 0SourceScholar
2020

Robust Low Rate Speech Coding Based on Cloned Networks and Wavenet

ICASSP 2020accepted

Rapid advances in machine-learning based generative modeling of speech make its use in speech coding attractive. However, the current performance of such models drops rapidly with noise contamination of the input, preventing use in practical applications. We present a new speech-coding scheme that i…

Cited by 0SourceScholar
2019

Differentiable Consistency Constraints for Improved Deep Speech Enhancement

ICASSP 2019accepted

In recent years, deep networks have led to dramatic improvements in speech enhancement by framing it as a data-driven pattern recognition problem. In many modern enhancement systems, large amounts of data are used to train a deep network to estimate masks for complex-valued short-time Fourier transf…

Cited by 0SourceScholar