← Search

Arvindh Krishnaswamy

8 accepted papers

2022

Improved Singing Voice Separation with Chromagram-Based Pitch-Aware Remixing

ICASSP 2022accepted

Singing voice separation aims to separate music into vocals and accompaniment components. One of the major constraints for the task is the limited amount of training data with separated vocals. Data augmentation techniques such as random source mixing have been shown to make better use of existing d…

Cited by 0SourceScholar
2022

Neural Speech Synthesis on a Shoestring: Improving the Efficiency of Lpcnet

ICASSP 2022accepted

Neural speech synthesis models can synthesize high quality speech but typically require a high computational complexity to do so. In previous work, we introduced LPCNet, which uses linear prediction to significantly reduce the complexity of neural synthesis. In this work, we further improve the effi…

Cited by 0SourceScholar
2021

Enhancing Audio Augmentation Methods with Consistency Learning

ICASSP 2021accepted

Data augmentation is an inexpensive way to increase training data diversity, and is commonly achieved via transformations of existing data. For tasks such as classification, there is a good case for learning representations of the data that are invariant to such transformations, yet this is not expl…

Cited by 0SourceScholar
2021

Enhancing into the Codec: Noise Robust Speech Coding with Vector-Quantized Autoencoders

ICASSP 2021accepted

Audio codecs based on discretized neural autoencoders have recently been developed and shown to provide significantly higher compression levels for comparable quality speech out-put. However, these models are tightly coupled with speech content, and produce unintended outputs in noisy conditions. Ba…

Cited by 0SourceScholar
2021

Low-Complexity, Real-Time Joint Neural Echo Control and Speech Enhancement Based On Percepnet

ICASSP 2021accepted

Speech enhancement algorithms based on deep learning have greatly surpassed their traditional counterparts and are now being considered for the task of removing acoustic echo from hands-free communication systems. This is a challenging problem due to both real-world constraints like loudspeaker non-…

Cited by 56SourceScholar
2021

Semi-Supervised Singing Voice Separation With Noisy Self-Training

ICASSP 2021accepted

Recent progress in singing voice separation has primarily focused on supervised deep learning methods. However, the scarcity of ground-truth data with clean musical sources has been a problem for long. Given a limited set of labeled data, we present a method to leverage a large volume of unlabeled d…

Cited by 0SourceScholar
2020

Channel-Attention Dense U-Net for Multichannel Speech Enhancement

ICASSP 2020accepted

Supervised deep learning has gained significant attention for speech enhancement recently. The state-of-the-art deep learning methods perform the task by learning a ratio/binary mask that is applied to the mixture in the time-frequency domain to produce the clean speech. Despite the great performanc…

Cited by 0SourceScholar
2020

Efficient Trainable Front-Ends for Neural Speech Enhancement

ICASSP 2020accepted

Many neural speech enhancement and source separation systems operate in the time-frequency domain. Such models often benefit from making their Short-Time Fourier Transform (STFT) front-ends trainable. In current literature, these are implemented as large Discrete Fourier Transform matrices; which ar…

Cited by 4SourceScholar