← Search

Rohan Kumar Das

14 accepted papers

2026

Environmental Sound Deepfake Detection Challenge: An Overview

ICASSP 2026poster

Recent progress in audio generation models has made it possible to create highly realistic and immersive soundscapes, which are now widely used in film and virtual-reality-related applications. However, these audio generators also raise concerns about potential misuse, such as producing deceptive au…

Cited by 0SourcePDFScholar
2026

Linking Faces and Voices Across Languages: Insights from the FAME 2026 Challenge

ICASSP 2026poster

Over half of the world's population is bilingual and people often communicate under multilingual scenarios. The Face-Voice Association in Multilingual Environments (FAME) 2026 Challenge, held at ICASSP 2026, focuses on developing methods for face-voice association that are effective when the languag…

Cited by 0SourcePDFScholar
2025

AnalyticKWS: Towards Exemplar-Free Analytic Class Incremental Learning for Small-footprint Keyword Spotting

ACL 2025finding

Keyword spotting (KWS) offers a vital mechanism to identify spoken commands in voice-enabled systems, where user demands often shift, requiring models to learn new keywords continually over time. However, a major problem is catastrophic forgetting, where models lose their ability to recognize earlie…

Cited by 0SourcePDFScholar
2025

Exploring Text-Queried Sound Event Detection with Audio Source Separation

ICASSP 2025accepted

In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we firs…

Cited by 0SourceScholar
2024

Adaptive-Avg-Pooling Based Attention Vision Transformer for Face Anti-Spoofing

ICASSP 2024accepted

Traditional vision transformer consists of two parts: transformer encoder and multi-layer perception (MLP). The former plays the role of feature learning to obtain better representation, while the latter plays the role of classification. Here, the MLP is constituted of two fully connected (FC) layer…

Cited by 0SourceScholar
2022

MFA: TDNN with Multi-Scale Frequency-Channel Attention for Text-Independent Speaker Verification with Short Utterances

ICASSP 2022accepted

The time delay neural network (TDNN) represents one of the state-of-the-art of neural solutions to text-independent speaker verification. However, they require a large number of filters to capture the speaker characteristics at any local frequency region. In addition, the performance of such systems…

Cited by 0SourceScholar
2022

Self-Supervised Speaker Recognition with Loss-Gated Learning

ICASSP 2022accepted

In self-supervised learning for speaker recognition, pseudo labels are useful as the supervision signals. It is a known fact that a speaker recognition model doesn’t always benefit from pseudo labels due to their unreliability. In this work, we observe that a speaker recognition network tends to mod…

Cited by 0SourceScholar
2021

Data Augmentation with Signal Companding for Detection of Logical Access Attacks

ICASSP 2021accepted

The recent advances in voice conversion (VC) and text-to-speech (TTS) make it possible to produce natural sounding speech that poses threat to automatic speaker verification (ASV) systems. To this end, research on spoofing countermeasures has gained attention to protect ASV systems from such attacks…

Cited by 0SourceScholar
2020

End-to-End Code-Switching TTS with Cross-Lingual Language Model

ICASSP 2020accepted

Code-switching text-to-speech (TTS) aims to enable a system to speak two languages with a single voice and in the same utterance. In this paper, we propose to incorporate cross-lingual word embedding into an end-to-end TTS system, to improve the voice rendering. The cross-lingual word embedding, gen…

Cited by 0SourceScholar
2020

On the Importance of Vocal Tract Constriction for Speaker Characterization: The Whispered Speech Study

ICASSP 2020accepted

Characterizing speakers under stressed condition is a challenge because speakers deviate from the normal speech production process. Whispered speech is one among them that is produced by abducting the vocal folds to pass the air out of mouth. During this process, the airflow thus passed is influence…

Cited by 0SourceScholar
2019

Cross-lingual Voice Conversion with Bilingual Phonetic Posteriorgram and Average Modeling

ICASSP 2019accepted

This paper presents a cross-lingual voice conversion approach using bilingual Phonetic PosteriorGram (PPG) and average modeling. The proposed approach makes use of bilingual PPGs to represent speaker-independent features of speech signals from different languages in the same feature space. In partic…

Cited by 0SourceScholar