← Search

Jan Skoglund

15 accepted papers

2025

Perceptual Audio Coding: A 40-Year Historical Perspective

ICASSP 2025accepted

In the history of audio and acoustic signal processing, perceptual audio coding has certainly excelled as a bright success story by its ubiquitous deployment in virtually all digital media devices, such as computers, tablets, mobile phones, set-top-boxes, and digital radios. From a technology perspe…

Cited by 0SourceScholar
2024

NOMAD: Unsupervised Learning of Perceptual Embeddings For Speech Enhancement and Non-Matching Reference Audio Quality Assessment

ICASSP 2024accepted

This paper presents NOMAD (Non-Matching Audio Distance), a differentiable perceptual similarity metric that measures the distance of a degraded signal against non-matching references. The proposed method is based on learning deep feature embeddings via a triplet loss guided by the Neurogram Similari…

Cited by 0SourceScholar
2024

SCOREQ: Speech Quality Assessment with Contrastive Regression

NeurIPS 2024poster

In this paper, we present SCOREQ, a novel approach for speech quality prediction. SCOREQ is a triplet loss function for contrastive regression that addresses the domain generalisation shortcoming exhibited by state of the art no-reference speech quality metrics. In the paper we: (i) illustrate the p…

2023

LMCodec: A Low Bitrate Speech Codec with Causal Transformer Models

ICASSP 2023accepted

We introduce LMCodec, a causal neural speech codec that provides high quality audio at very low bitrates. The backbone of the system is a causal convolutional codec that encodes audio into a hierarchy of coarse-to-fine tokens using residual vector quantization. LMCodec trains a Transformer language…

Cited by 0SourceScholar
2021

Generative Speech Coding with Predictive Variance Regularization

ICASSP 2021accepted

The recent emergence of machine-learning based generative models for speech suggests a significant reduction in bit rate for speech codecs is possible. However, the performance of generative models deteriorates significantly with the distortions present in real-world input signals. We argue that thi…

Cited by 0SourceScholar
2021

Warp-Q: Quality Prediction for Generative Neural Speech Codecs

ICASSP 2021accepted

Good speech quality has been achieved using waveform matching and parametric reconstruction coders. Recently developed very low bit rate generative codecs can reconstruct high quality wideband speech with bit streams less than 3 kb/s. These codecs use a DNN with parametric input to synthesise high q…

Cited by 0SourceScholar
2020

Robust Low Rate Speech Coding Based on Cloned Networks and Wavenet

ICASSP 2020accepted

Rapid advances in machine-learning based generative modeling of speech make its use in speech coding attractive. However, the current performance of such models drops rapidly with noise contamination of the input, preventing use in practical applications. We present a new speech-coding scheme that i…

Cited by 0SourceScholar
2018

Wavenet Based Low Rate Speech Coding

ICASSP 2018accepted

Traditional parametric coding of speech facilitates low rate but provides poor reconstruction quality because of the inadequacy of the model used. We describe how a WaveNet generative speech model can be used to generate high quality speech from the bit stream of a standard parametric coder operatin…

Cited by 155SourceScholar
2017

Practically efficient nonlinear acoustic echo cancellers using cascaded block RLS and FLMS adaptive filters

ICASSP 2017accepted

This paper presents a practically efficient implementation for non-linear acoustic echo cancellation (NAEC). The echo path is modeled by a novel hybrid Taylor-Volterra pre-processor followed by a linear FIR filter. A cascaded block RLS and unconstrained FLMS adaptive algorithm is developed to jointl…

Cited by 21SourceScholar
2016

An acoustic keystroke transient canceler for speech communication terminals using a semi-blind adaptive filter model

ICASSP 2016accepted

In many teleconferencing applications using modern laptop and net-book devices it is common to encounter annoying keyboard typing noise. In this paper we propose an acoustic keystroke transient canceler for speech communication terminals as a novel broadband adaptive filter application in such a han…

Cited by 0SourceScholar
2016

Globally optimized least-squares post-filtering for microphone array speech enhancement

ICASSP 2016accepted

Existing post-filtering techniques for microphone array speech enhancement have two common deficiencies. First, they assume that the noise is either white or diffuse and cannot deal with point inter-ferers. Second, they estimate the post-filter coefficients using only two microphones at a time and t…

Cited by 21SourceScholar
2015

Detection and suppression of keyboard transient noise in audio streams with auxiliary keybed microphone

ICASSP 2015accepted

In this paper a problem in transient noise suppression for audio streams in laptop and netbook devices is addressed. One or more microphones record voice signals which are corrupted with ambient noise and also transient noise from keyboard and mouse clicks. In the current work, a synchronous referen…

Cited by 0SourceScholar
2015

Direct-to-Reverberant Ratio estimation using a null-steered beamformer

ICASSP 2015accepted

Reverberation affects the quality and intelligibility of distant speech recorded in a room. Direct-to-Reverberant Ratio (DRR) is a useful measure for assessing the acoustic configuration and can be used to inform dereverberation algorithms. We describe a novel DRR estimation algorithm applicable whe…

Cited by 0SourceScholar