← Search

Ritwik Giri

12 accepted papers

2024

Real-Time Stereo Speech Enhancement with Spatial-Cue Preservation Based on Dual-Path Structure

ICASSP 2024accepted

We introduce a real-time, multichannel speech enhancement algorithm which maintains the spatial cues of stereo recordings including two speech sources. Recognizing that each source has unique spatial information, our method utilizes a dual-path structure, ensuring the spatial cues remain unaffected…

Cited by 0SourceScholar
2023

A Framework for Unified Real-Time Personalized and Non-Personalized Speech Enhancement

ICASSP 2023accepted

In this study, we present an approach to train a single speech enhancement network that can perform both personalized and non-personalized speech enhancement. This is achieved by incorporating a frame-wise conditioning input that specifies the type of enhancement output. To improve the quality of th…

Cited by 0SourceScholar
2022

Improved Singing Voice Separation with Chromagram-Based Pitch-Aware Remixing

ICASSP 2022accepted

Singing voice separation aims to separate music into vocals and accompaniment components. One of the major constraints for the task is the limited amount of training data with separated vocals. Data augmentation techniques such as random source mixing have been shown to make better use of existing d…

Cited by 0SourceScholar
2021

Enhancing into the Codec: Noise Robust Speech Coding with Vector-Quantized Autoencoders

ICASSP 2021accepted

Audio codecs based on discretized neural autoencoders have recently been developed and shown to provide significantly higher compression levels for comparable quality speech out-put. However, these models are tightly coupled with speech content, and produce unintended outputs in noisy conditions. Ba…

Cited by 0SourceScholar
2021

Semi-Supervised Singing Voice Separation With Noisy Self-Training

ICASSP 2021accepted

Recent progress in singing voice separation has primarily focused on supervised deep learning methods. However, the scarcity of ground-truth data with clean musical sources has been a problem for long. Given a limited set of labeled data, we present a method to leverage a large volume of unlabeled d…

Cited by 0SourceScholar
2020

Channel-Attention Dense U-Net for Multichannel Speech Enhancement

ICASSP 2020accepted

Supervised deep learning has gained significant attention for speech enhancement recently. The state-of-the-art deep learning methods perform the task by learning a ratio/binary mask that is applied to the mixture in the time-frequency domain to produce the clean speech. Despite the great performanc…

Cited by 0SourceScholar
2018

Improved Noise Characterization for Relative Impulse Response Estimation

ICASSP 2018accepted

Relative Impulse Responses (ReIRs) have several applications in speech enhancement, noise suppression and source localization for multi-channel speech processing in reverberant environments. Noise is usually assumed to be white Gaussian during the estimation of the ReIR between two microphones. We s…

Cited by 0SourceScholar
2016

Dynamic relative impulse response estimation using structured sparse Bayesian learning

ICASSP 2016accepted

In this paper we present a novel Hierarchical Bayesian approach to estimate Relative Impulse Response (ReIR) using short, noisy and reverberant microphone recordings. The information contained in ReIRs between two microphones is useful for a wide range of multichannel speech processing applications…

Cited by 10SourceScholar
2015

Improving speech recognition in reverberation using a room-aware deep neural network and multi-task learning

ICASSP 2015accepted

In this paper, we propose two approaches to improve deep neural network (DNN) acoustic models for speech recognition in reverberant environments. Both methods utilize auxiliary information in training the DNN but differ in the type of information and the manner in which it is used. The first method…

Cited by 0SourceScholar