← Search

Guillaume Fuchs

12 accepted papers

2025

A first-order DirAC-based parametric Ambisonic coder for immersive communications

ICASSP 2025accepted

Directional Audio Coding (DirAC) is a proven method for parametrically representing a 3D audio scene in B-format and is capable of reproducing it on arbitrary loudspeaker layouts. Although such a method seems well suited for low bitrate Ambisonic transmission, little work has been done on the feasib…

Cited by 0SourceScholar
2025

Ambisonics Coding in IVAS: A Hybrid SPAR and DirAC System

ICASSP 2025accepted

The Ambisonics audio format represents a 3D sound field as a fixed set of audio channels, with a tradeoff between spatial detail and the required number of audio channels. Because the number of audio channels that must be coded increases quadratically with respect to Ambisonics order, high-quality c…

Cited by 0SourceScholar
2025

Hybrid predictive and parametric stereo coding for voice and audio communications

ICASSP 2025accepted

The transmission of stereo audio in voice calls helps to improve immersion and user experience. However stereo coding has primarily been studied for broadcast or streaming applications. To extend the capabilities of 3GPP EVS, a hybrid stereo coding scheme combining predictive and parametric coding i…

Cited by 0SourceScholar
2025

Parametric Object Coding in IVAS: Efficient Coding of Multiple Audio Objects at Low Bit Rates

ICASSP 2025accepted

The recently standardized 3GPP codec for Immersive Voice and Audio Services (IVAS) includes a parametric mode for efficiently coding multiple audio objects at low bit rates. In this mode, parametric side information is obtained from both the object metadata and the input audio objects. The side info…

Cited by 0SourceScholar
2022

A DNN Based Post-Filter to Enhance the Quality of Coded Speech in MDCT Domain

ICASSP 2022accepted

Frequency domain processing, and in particular the use of Modified Discrete Cosine Transform (MDCT), is the most widespread approach to audio coding. However, at low bitrates, audio quality, especially for speech, degrades drastically due to the lack of available bits to directly code the transform…

Cited by 0SourceScholar
2022

PostGAN: A GAN-Based Post-Processor to Enhance the Quality of Coded Speech

ICASSP 2022accepted

The quality of speech coded by transform coding is affected by various artefacts especially when bitrates to quantize the frequency components become too low. In order to mitigate these coding artefacts and enhance the quality of coded speech, a post-processor that relies on a-priori information tra…

Cited by 0SourceScholar
2021

StyleMelGAN: An Efficient High-Fidelity Adversarial Vocoder with Temporal Adaptive Normalization

ICASSP 2021accepted

In recent years, neural vocoders have surpassed classical speech generation approaches in naturalness and perceptual quality of the synthesized speech. Computationally heavy models like WaveNet and WaveGlow achieve best results, while lightweight GAN models, e.g. MelGAN and Parallel WaveGAN, remain…

Cited by 0SourceScholar
2018

GMM-Based Iterative Entropy Coding for Spectral Envelopes of Speech and Audio

ICASSP 2018accepted

Spectral envelope modelling is a central part of speech and audio codecs and is traditionally based on either vector quantization or scalar quantization followed by entropy coding. To bridge the coding performance of vector quantization with the low complexity of the scalar case, we propose an itera…

Cited by 0SourceScholar
2015

Frequency-domain Comfort Noise Generation for Discontinuous Transmission in EVS

ICASSP 2015accepted

Discontinuous Transmission (DTX) is an efficient way to drastically reduce the transmission rate of a communication codec in the absence of voice input. In this mode, most frames that are determined to consist of background noise only are dropped from transmission and replaced by some Comfort Noise…

Cited by 0SourceScholar
2015

Low delay LPC and MDCT-based audio coding in the EVS codec

ICASSP 2015accepted

Speech coders operating in time domain can be extended with a frequency domain mode to improve encoding of music, even though this is challenging at low delay. In such a scenario, the short analysis window limits the benefit of the transform coder, while a delayless switch between the two coders con…

Cited by 0SourceScholar
2015

Low-complexity and robust coding mode decision in the EVS coder

ICASSP 2015accepted

Several state-of-the-art switched audio codecs employ the closed-loop mode decision to select the best coding mode at every frame. The closed-loop mode selection is known to have good performance but also high complexity. The new approach we propose in this paper is a low-complexity version of the c…

Cited by 0SourceScholar