← Search

Jürgen Herre

6 accepted papers

2026

DEEPAQ: A PERCEPTUAL AUDIO QUALITY METRIC BASED ON FOUNDATIONAL MODELS AND WEAKLY SUPERVISED LEARNING

ICASSP 2026oral

This paper presents the Deep learning-based Perceptual Audio Quality metric (DeePAQ) for evaluating general audio quality. Our approach leverages metric learning together with the music foundation model MERT, guided by surrogate labels, to construct an embedding space that captures distortion intens…

Cited by 0SourcePDFScholar
2025

Perceptual Audio Coding: A 40-Year Historical Perspective

ICASSP 2025accepted

In the history of audio and acoustic signal processing, perceptual audio coding has certainly excelled as a bright success story by its ubiquitous deployment in virtually all digital media devices, such as computers, tablets, mobile phones, set-top-boxes, and digital radios. From a technology perspe…

Cited by 0SourceScholar
2022

A Data-Driven Cognitive Salience Model for Objective Perceptual Audio Quality Assessment

ICASSP 2022accepted

Objective audio quality measurement systems often use perceptual models to predict the subjective quality scores of processed signals, as reported in listening tests. Most systems map different metrics of perceived degradation into a single quality score predicting subjective quality. This requires…

Cited by 0SourceScholar
2019

Predicting the Precision of Elevation Localization Based on Head Related Transfer Functions

ICASSP 2019accepted

While the human hearing capability for horizontally localized sound sources is mostly based on binaural cues, the localization of elevation along the "cones of confusion" can only rely on spectral cues provided by the directionality of head related transfer functions (HRTF). This paper explores how…

Cited by 0SourceScholar
2017

Coding of fine granular audio signals using High Resolution Envelope Processing (HREP)

ICASSP 2017accepted

High Resolution Envelope Processing (HREP) is a new tool for improved perceptual coding of audio signals that predominantly consist of many dense transient events, such as applause, rain drop sounds, etc. These signals have traditionally been very difficult to code for perceptual audio codecs, parti…

Cited by 0SourceScholar