← Search

Jean-Marc Valin

13 accepted papers

2024

NOLACE: Improving Low-Complexity Speech Codec Enhancement Through Adaptive Temporal Shaping

ICASSP 2024accepted

Speech codec enhancement methods are designed to remove distortions added by speech codecs. While classical methods are very low in complexity and add zero delay, their effectiveness is rather limited. Compared to that, DNN-based methods deliver higher quality but they are typically high in complexi…

Cited by 0SourceScholar
2024

Noise-Robust DSP-Assisted Neural Pitch Estimation With Very Low Complexity

ICASSP 2024accepted

Pitch estimation is an essential step of many speech processing algorithms, including speech coding, synthesis, and enhancement. Recently, pitch estimators based on deep neural networks (DNNs) have been outperforming well-established DSP-based techniques. Unfortunately, these new estimators can be i…

Cited by 0SourceScholar
2024

Real-Time Stereo Speech Enhancement with Spatial-Cue Preservation Based on Dual-Path Structure

ICASSP 2024accepted

We introduce a real-time, multichannel speech enhancement algorithm which maintains the spatial cues of stereo recordings including two speech sources. Recognizing that each source has unique spatial information, our method utilizes a dual-path structure, ensuring the spatial cues remain unaffected…

Cited by 0SourceScholar
2023

A Framework for Unified Real-Time Personalized and Non-Personalized Speech Enhancement

ICASSP 2023accepted

In this study, we present an approach to train a single speech enhancement network that can perform both personalized and non-personalized speech enhancement. This is achieved by incorporating a frame-wise conditioning input that specifies the type of enhancement output. To improve the quality of th…

Cited by 10SourceScholar
2023

Framewise Wavegan: High Speed Adversarial Vocoder In Time Domain With Very Low Computational Complexity

ICASSP 2023accepted

GAN vocoders are currently one of the state-of-the-art methods for building high-quality neural waveform generative models. However, most of their architectures require dozens of billion floating-point operations per second (GFLOPS) to generate speech waveforms in samplewise manner. This makes GAN v…

Cited by 8SourceScholar
2023

Low-Bitrate Redundancy Coding of Speech Using A Rate-Distortion-Optimized Variational Autoencoder

ICASSP 2023accepted

Robustness to packet loss is one of the main ongoing challenges in real-time speech communication. Deep packet loss concealment (PLC) techniques have recently demonstrated improved quality compared to traditional PLC. Despite that, all PLC techniques hit fundamental limitations when too much acousti…

Cited by 0SourceScholar
2022

Improved Singing Voice Separation with Chromagram-Based Pitch-Aware Remixing

ICASSP 2022accepted

Singing voice separation aims to separate music into vocals and accompaniment components. One of the major constraints for the task is the limited amount of training data with separated vocals. Data augmentation techniques such as random source mixing have been shown to make better use of existing d…

Cited by 0SourceScholar
2022

Neural Speech Synthesis on a Shoestring: Improving the Efficiency of Lpcnet

ICASSP 2022accepted

Neural speech synthesis models can synthesize high quality speech but typically require a high computational complexity to do so. In previous work, we introduced LPCNet, which uses linear prediction to significantly reduce the complexity of neural synthesis. In this work, we further improve the effi…

Cited by 0SourceScholar
2021

Enhancing into the Codec: Noise Robust Speech Coding with Vector-Quantized Autoencoders

ICASSP 2021accepted

Audio codecs based on discretized neural autoencoders have recently been developed and shown to provide significantly higher compression levels for comparable quality speech out-put. However, these models are tightly coupled with speech content, and produce unintended outputs in noisy conditions. Ba…

Cited by 0SourceScholar
2021

Low-Complexity, Real-Time Joint Neural Echo Control and Speech Enhancement Based On Percepnet

ICASSP 2021accepted

Speech enhancement algorithms based on deep learning have greatly surpassed their traditional counterparts and are now being considered for the task of removing acoustic echo from hands-free communication systems. This is a challenging problem due to both real-world constraints like loudspeaker non-…

Cited by 56SourceScholar
2021

Semi-Supervised Singing Voice Separation With Noisy Self-Training

ICASSP 2021accepted

Recent progress in singing voice separation has primarily focused on supervised deep learning methods. However, the scarcity of ground-truth data with clean musical sources has been a problem for long. Given a limited set of labeled data, we present a method to leverage a large volume of unlabeled d…

Cited by 0SourceScholar