← Search

Karan Thakkar

4 accepted papers

2025

SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer

ICASSP 2025accepted

In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer that operates on latent features. SoloAudio supports both a…

Cited by 0SourceScholar
2024

DPM-TSE: A Diffusion Probabilistic Model for Target Sound Extraction

ICASSP 2024accepted

Common target sound extraction (TSE) approaches primarily relied on discriminative approaches in order to separate the target sound while minimizing interference from the unwanted sources, with varying success in separating the target from the background. This study introduces DPM-TSE, a generative…

Cited by 0SourceScholar
2024

Investigating Self-Supervised Deep Representations for EEG-Based Auditory Attention Decoding

ICASSP 2024accepted

Auditory Attention Decoding (AAD) algorithms play a crucial role in isolating desired sound sources within challenging acoustic environments directly from brain activity. Although recent research has shown promise in AAD using shallow representations such as auditory envelope and spectrogram, there…

Cited by 0SourceScholar