← Search

Ankur Kumar

3 accepted papers

2025

What Does an Audio Deepfake Detector Focus on? A Study in the Time Domain

ICASSP 2025accepted

Adding explanations to audio deepfake detection (ADD) models will enable insights on the decision making process and thus boost their real-world application. In this paper, we propose a relevancy-based explainable AI (XAI) method to analyze the predictions of transformer-based ADD models. We compare…

Cited by 0SourceScholar
2024

Stateful Conformer with Cache-Based Inference for Streaming Automatic Speech Recognition

ICASSP 2024accepted

In this paper, we propose an efficient and accurate streaming speech recognition model based on the FastConformer architecture. We adapted the FastConformer architecture for streaming applications through: (1) constraining both the look-ahead and past contexts in the encoder, and (2) introducing an…

Cited by 0SourceScholar
2021

Neural Utterance Confidence Measure for RNN-Transducers and Two Pass Models

ICASSP 2021accepted

In this paper, we propose methods to compute confidence score on the predictions made by an end-to-end speech recognition model in a 2-pass framework. We use RNN-Transducer for a streaming model, and an attention-based decoder for the second pass model. We use neural technique to compute the confide…

Cited by 0SourceScholar