← Search

Vesa Välimäki

14 accepted papers

2026

ANYRIR: ROBUST NON-INTRUSIVE ROOM IMPULSE RESPONSE ESTIMATION IN THE WILD

ICASSP 2026poster

We address the problem of estimating room impulse responses (RIRs) in noisy, uncontrolled environments where non-stationary sounds such as speech or footsteps corrupt conventional deconvolution. We propose AnyRIR, a non-intrusive method that uses music as the excitation signal instead of a dedicated…

Cited by 0SourcePDFScholar
2026

Automatic Music Mixing using a Generative Model of Effect Embeddings

ICASSP 2026oral

Music mixing involves combining individual tracks into a cohesive mixture, a task characterized by subjectivity where multiple valid solutions exist for the same input. Existing automatic mixing systems treat this task as a deterministic regression problem, thus ignoring this multiplicity of solutio…

Cited by 0SourcePDFScholar
2026

MATCHING REVERBERANT SPEECH THROUGH LEARNED ACOUSTIC EMBEDDINGS AND FEEDBACK DELAY NETWORKS

ICASSP 2026oral

Reverberation conveys critical acoustic cues about the environment, supporting spatial awareness and immersion. For auditory augmented reality (AAR) systems, generating perceptually plausible reverberation in real time remains a key challenge, especially when explicit acoustic measurements are unava…

Cited by 0SourcePDFScholar
2025

FLAMO: An Open-Source Library for Frequency-Domain Differentiable Audio Processing

ICASSP 2025accepted

We present FLAMO, a Frequency-sampling Library for Audio-Module Optimization designed to implement and optimize differentiable linear time-invariant audio systems. The library is open-source and built on the frequency-sampling filter design method, allowing for the creation of differentiable modules…

Cited by 0SourceScholar
2025

HRTF Estimation using a Score-based Prior

ICASSP 2025accepted

We present a head-related transfer function (HRTF) estimation method which relies on a data-driven prior given by a score-based diffusion model. The HRTF is estimated in reverberant environments using natural excitation signals, e.g. human speech. The impulse response of the room is estimated along…

Cited by 0SourceScholar
2023

Extreme Audio Time Stretching Using Neural Synthesis

ICASSP 2023accepted

A deep neural network solution for time-scale modification (TSM) focused on large stretching factors is proposed, targeting environmental sounds. Traditional TSM artifacts such as transient smearing, loss of presence, and phasiness are heavily accentuated and cause poor audio quality when the TSM fa…

Cited by 0SourceScholar
2022

Audio Peak Reduction Using a Synced allpass Filter

ICASSP 2022accepted

Peak reduction is a common step used in audio playback chains to increase the loudness of a sound. The distortion introduced by a conventional nonlinear compressor can be avoided with the use of an allpass filter, which provides peak reduction by acting on the signal phase. This way, the signal ener…

Cited by 7SourceScholar
2019

Graphic Delay Equalizer

ICASSP 2019accepted

A graphic delay equalizer based on a high-order nonparametric allpass filter design is proposed. Command points at the centers of octave frequency bands are connected with polynomial interpolation to form a continuous target group-delay curve as function of frequency. The required number of all-pass…

Cited by 4SourceScholar