← Search

Ante Jukic

7 accepted papers

2026

GENERALIZABILITY OF PREDICTIVE AND GENERATIVE SPEECH ENHANCEMENT MODELS TO PATHOLOGICAL SPEAKERS

ICASSP 2026poster

State of the art speech enhancement (SE) models achieve strong performance on neurotypical speech, but their effectiveness is substantially reduced for pathological speech. In this paper, we investigate strategies to address this gap for both predictive and generative SE models, including i) trainin…

Cited by 0SourcePDFScholar
2025

Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration

ICASSP 2025accepted

This paper proposes a generative pretraining foundation model for high-quality speech restoration tasks. By directly operating on complex-valued short-time Fourier transform coefficients, our model does not rely on any vocoders for time-domain signal reconstruction. As a result, our model simplifies…

Cited by 0SourceScholar
2025

Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference

ICASSP 2025accepted

Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often operate at high frame rates, resulting in slow training and infe…

Cited by 0SourceScholar
2025

Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement

ICASSP 2025accepted

In this work, we investigate application of generative speech enhancement to improve the robustness of ASR models in noisy and reverberant conditions. We employ a recently-proposed speech enhancement model based on Schrödinger bridge, which has been shown to perform well compared to diffusion-based…

Cited by 0SourceScholar
2016

Robust sparsity-promoting acoustic multi-channel equalization for speech dereverberation

ICASSP 2016accepted

This paper presents a novel signal-dependent method to increase the robustness of acoustic multi-channel equalization techniques against room impulse response (RIR) estimation errors. Aiming at obtaining an output signal which better resembles a clean speech signal, we propose to extend the acoustic…

Cited by 0SourceScholar
2015

Multi-channel linear prediction-based speech dereverberation with low-rank power spectrogram approximation

ICASSP 2015accepted

In many acoustic conditions the recorded speech signals may be severely affected by reverberation, leading to a reduced speech quality and intelligibility. In this paper we focus on a blind speech dereverberation method based on multi-channel linear prediction (MCLP) in the short-time Fourier transf…

Cited by 0SourceScholar