← Search

Paris Smaragdis

28 accepted papers

2026

GENCHO: ROOM IMPULSE RESPONSE GENERATION FROM REVERBERANT SPEECH AND TEXT VIA DIFFUSION TRANSFORMERS

ICASSP 2026oral

Blind room impulse response (RIR) estimation is a core task for capturing and transferring acoustic properties; yet existing methods often suffer from limited modeling capability and degraded performance under unseen conditions. Moreover, emerging generative audio applications call for more flexible…

Cited by 0SourcePDFScholar
2026

PROMPTSEP: GENERATIVE AUDIO SEPARATION VIA MULTIMODAL PROMPTING

ICASSP 2026oral

Recent breakthroughs in language-queried audio source separation (LASS) have shown that generative models can achieve higher separation audio quality than traditional masking-based approaches. However, two key limitations restrict their practical use: (1) users often require operations beyond separa…

Cited by 0SourcePDFScholar
2025

On Class Separability Pitfalls In Audio-Text Contrastive Zero-Shot Learning

ICASSP 2025accepted

Recent advances in audio-text cross-modal contrastive learning have shown its potential towards zero-shot learning. One possibility for this is by projecting item embeddings from pre-trained backbone neural networks into a cross-modal space in which item similarity can be calculated in either domain…

Cited by 0SourceScholar
2024

Noise-Robust DSP-Assisted Neural Pitch Estimation With Very Low Complexity

ICASSP 2024accepted

Pitch estimation is an essential step of many speech processing algorithms, including speech coding, synthesis, and enhancement. Recently, pitch estimators based on deep neural networks (DNNs) have been outperforming well-established DSP-based techniques. Unfortunately, these new estimators can be i…

Cited by 0SourceScholar
2023

A Framework for Unified Real-Time Personalized and Non-Personalized Speech Enhancement

ICASSP 2023accepted

In this study, we present an approach to train a single speech enhancement network that can perform both personalized and non-personalized speech enhancement. This is achieved by incorporating a frame-wise conditioning input that specifies the type of enhancement output. To improve the quality of th…

Cited by 0SourceScholar
2023

Framewise Wavegan: High Speed Adversarial Vocoder In Time Domain With Very Low Computational Complexity

ICASSP 2023accepted

GAN vocoders are currently one of the state-of-the-art methods for building high-quality neural waveform generative models. However, most of their architectures require dozens of billion floating-point operations per second (GFLOPS) to generate speech waveforms in samplewise manner. This makes GAN v…

Cited by 0SourceScholar
2023

Generative Modeling Based Manifold Learning for Adaptive Filtering Guidance

ICASSP 2023accepted

In most practical adaptive filtering problems, estimated filters are not arbitrary, but instead lie on a manifold that encapsulates characteristics of the problem at hand. Consequently, it is desirable to steer adaptation towards filters that lie on that manifold. In this paper, we propose a novel a…

Cited by 0SourceScholar
2023

Latent Iterative Refinement for Modular Source Separation

ICASSP 2023accepted

Traditional source separation approaches train deep neural network models end-to-end with all the data available at once by minimizing the empirical risk on the whole training set. On the inference side, after training the model, the user fetches a static computation graph and runs the full model on…

Cited by 0SourceScholar
2023

Optimal Condition Training for Target Source Separation

ICASSP 2023accepted

Recent research has shown remarkable performance in leveraging multiple extraneous conditional and non-mutually-exclusive semantic concepts for sound source separation, allowing the flexibility to extract a given target source based on multiple different queries. In this work, we propose a new optim…

Cited by 0SourceScholar
2022

Neural Speech Synthesis on a Shoestring: Improving the Efficiency of Lpcnet

ICASSP 2022accepted

Neural speech synthesis models can synthesize high quality speech but typically require a high computational complexity to do so. In previous work, we introduced LPCNet, which uses linear prediction to significantly reduce the complexity of neural synthesis. In this work, we further improve the effi…

Cited by 0SourceScholar
2021

Communication-Cost Aware Microphone Selection for Neural Speech Enhancement with Ad-Hoc Microphone Arrays

ICASSP 2021accepted

In this paper, we present a method for jointly-learning a microphone selection mechanism and a speech enhancement network for multi-channel speech enhancement with an ad-hoc microphone array. The attention-based microphone selection mechanism is trained to reduce communication-costs through a penalt…

Cited by 0SourceScholar
2021

Differentiable Signal Processing With Black-Box Audio Effects

ICASSP 2021accepted

We present a data-driven approach to automate audio signal processing by incorporating stateful third-party, audio effects as layers within a deep neural network. We then train a deep encoder to analyze input audio and control effect parameters to perform the desired signal manipulation, requiring o…

Cited by 0SourceScholar
2021

Unified Gradient Reweighting for Model Biasing with Applications to Source Separation

ICASSP 2021accepted

Recent deep learning approaches have shown great improvement in audio source separation tasks. However, the vast majority of such work is focused on improving average separation performance, often neglecting to examine or control the distribution of the results. In this paper, we propose a simple, u…

Cited by 0SourceScholar
2020

End-To-End Non-Negative Autoencoders for Sound Source Separation

ICASSP 2020accepted

Discriminative models for source separation have recently been shown to produce impressive results. However, when operating on sources outside of the training set, these models can not perform as well and are cumbersome to update. Classical methods like Nonnegative Matrix Factorization (NMF) provide…

Cited by 0SourceScholar
2020

One-Shot Parametric Audio Production Style Transfer with Application to Frequency Equalization

ICASSP 2020accepted

Audio production is a difficult process for many people], [and properly manipulating sound to achieve a certain effect is non-trivial. In this paper], [we present a method that facilitates this process by inferring appropriate audio effect parameters in order to make an input recording sound similar…

Cited by 0SourceScholar
2020

Two-Step Sound Source Separation: Training On Learned Latent Targets

ICASSP 2020accepted

In this paper, we propose a two-step training procedure for source separation via a deep neural network. In the first step we learn a transform (and it's inverse) to a latent space where masking-based separation performance using oracles is optimal. For the second step, we train a separation module…

Cited by 0SourceScholar
2019

Majorization-minimization Algorithms for Convolutive NMF with the Beta-divergence

ICASSP 2019accepted

Nonnegative matrix factorization (NMF) has become a method of choice for spectrogram decomposition. However, its inability to capture dependencies across columns of the input motivated the introduction of a variant, convolutive NMF. While algorithms for solving the convolutive NMF problem were previ…

Cited by 0SourceScholar
2019

Unsupervised Deep Clustering for Source Separation: Direct Learning from Mixtures Using Spatial Information

ICASSP 2019accepted

We present a monophonic source separation system that is trained by only observing mixtures with no ground truth separation information. We use a deep clustering approach which trains on multichannel mixtures and learns to project spectrogram bins to source clusters that correlate with various spati…

Cited by 0SourceScholar
2018

Blind Estimation of the Speech Transmission Index for Speech Quality Prediction

ICASSP 2018accepted

The speech transmission index (STI) of a listening position within a given room indicates the quality and intelligibility of speech uttered in that room. The measure is very reliable for predicting speech intelligibility in many room conditions but requires an STI measurement of the impulse response…

Cited by 0SourceScholar
2016

Efficient neighborhood-based topic modeling for collaborative audio enhancement on massive crowdsourced recordings

ICASSP 2016accepted

Collaborative Audio Enhancement (CAE) aims at separating a dominant source from crowdsourced recordings of a scene. This paper proposes a CAE setup as a big ad-hoc microphone array problem, assuming hundreds of sensors scattered over a large scene, e.g. a concert hall or a street riot. An important…

Cited by 0SourceScholar
2015

Efficient manifold preserving audio source separation using locality sensitive hashing

ICASSP 2015accepted

We propose an efficient technique to learn probabilistic hierarchical topic models that are designed to preserve the manifold structure of audio data. The consideration of the data manifold is important, as it has been shown to provide superior performance in certain audio applications such as sourc…

Cited by 0SourceScholar
2015

Joint acoustic and spectral modeling for speech dereverberation using non-negative representations

ICASSP 2015accepted

This paper proposes a single-channel speech dereverberation method enhancing the spectrum of the reverberant speech signal. The proposed method uses a non-negative approximation of the convolutive transfer function (N-CTF) to simultaneously estimate the magnitude spectrograms of the speech signal an…

Cited by 0SourceScholar