← Search

Sriram Ganapathy

33 accepted papers

2025

FESTA: Functionally Equivalent Sampling for Trust Assessment of Multimodal LLMs

EMNLP 2025

The accurate trust assessment of multimodal large language models (MLLMs) generated predictions, which can enable selective prediction and improve user confidence, is challenging due to the diverse multi-modal input paradigms. We propose F unctionally E quivalent S ampling for T rust A ssessment (FE

2025

Identifying and Mitigating Mismatched Language Code in Multilingual ASR

ICASSP 2025accepted

Multilingual speech recognition systems often use an input language code in order to prompt the transcription in the target language. However, the spoken language in the input audio may not always match the language code, as often prevalent in multilingual societies. This language mismatch can signi…

Cited by 0SourceScholar
2024

LLM Augmented LLMs: Expanding Capabilities through Composition

ICLR 2024poster

Foundational models with billions of parameters which have been trained on large corpus of data have demonstrated non-trivial skills in a variety of domains. However, due to their monolithic structure, it is challenging and expensive to augment them or impart new skills. On the other hand, due to th…

Cited by 44SourcePDFScholar
2024

Multimodal Modeling for Spoken Language Identification

ICASSP 2024accepted

Spoken language identification refers to the task of automatically predicting the spoken language in a given utterance. Conventionally, it is modeled as a speech-based language identification task. Prior techniques have been constrained to a single modality; however in the case of video data there i…

Cited by 0SourceScholar
2023

Accented Speech Recognition With Accent-specific Codebooks

EMNLP 2023long main

Speech accents pose a significant challenge to state-of-the-art automatic speech recognition (ASR) systems. Degradation in performance across underrepresented accents is a severe deterrent to the inclusive adoption of ASR. In this work, we propose a novel accent adaptation approach for end-to-end AS…

Cited by 0SourcecodeScholar
2023

Self-Influence Guided Data Reweighting for Language Model Pre-training

EMNLP 2023long main

Language Models (LMs) pre-trained with selfsupervision on large text corpora have become the default starting point for developing models for various NLP tasks. Once the pre-training corpus has been assembled, all data samples in the corpus are treated with equal importance during LM pre-training. H…

Cited by 0SourceScholar
2023

Supervised Hierarchical Clustering Using Graph Neural Networks for Speaker Diarization

ICASSP 2023accepted

Conventional methods for speaker diarization involve windowing an audio file into short segments to extract speaker embeddings, followed by an unsupervised clustering of the embeddings. This multistep approach generates speaker assignments for each segment. In this paper, we propose a novel Supervis…

Cited by 0SourceScholar
2022

End-To-End Speech Recognition with Joint Dereverberation of Sub-Band Autoregressive Envelopes

ICASSP 2022accepted

The end-to-end (E2E) automatic speech recognition (ASR) systems are often required to operate in reverberant conditions, where the long-term sub-band envelopes of the speech are temporally smeared. In this paper, we develop a feature enhancement approach using a neural model operating on sub-band te…

Cited by 0SourceScholar
2022

Multimodal Transformer with Learnable Frontend and Self Attention for Emotion Recognition

ICASSP 2022accepted

In this work, we propose a novel approach for multi-modal emotion recognition from conversations using speech and text. The audio representations are learned jointly with a learnable audio front-end (LEAF) model feeding to a CNN based classifier. The text representations are derived from pre-trained…

Cited by 0SourceScholar
2022

Self Supervised Representation Learning with Deep Clustering for Acoustic Unit Discovery from Raw Speech

ICASSP 2022accepted

The automatic discovery of acoustic sub-word units from raw speech, without any text or labels, is a growing field of research. The key challenge is to derive representations of speech that can be categorized into a small number of phoneme-like units which are speaker invariant and can broadly captu…

Cited by 0SourceScholar
2022

The Second Dicova Challenge: Dataset and Performance Analysis for Diagnosis of Covid-19 Using Acoustics

ICASSP 2022accepted

The Second Diagnosis of COVID-19 using Acoustics (DiCOVA) Challenge aimed at accelerating the research in acoustics based detection of COVID-19, a topic at the intersection of acoustics, signal processing, machine learning, and healthcare. This paper presents the details of the challenge, which was…

Cited by 0SourceScholar
2021

Deep Multiway Canonical Correlation Analysis For Multi-Subject Eeg Normalization

ICASSP 2021accepted

The normalization of brain recordings from multiple subjects responding to the natural stimuli is one of the key challenges in auditory neuroscience. The objective of this normalization is to transform the brain data in such a way as to remove the inter-subject redundancies and to boost the componen…

Cited by 0SourceScholar
2021

End-to-End Lyrics Recognition with Voice to Singing Style Transfer

ICASSP 2021accepted

Automatic transcription of monophonic/polyphonic music is a challenging task due to the lack of availability of large amounts of transcribed data. In this paper, we propose a data augmentation method that converts natural speech to singing voice based on vocoder based speech synthesizer. This approa…

Cited by 0SourceScholar
2021

NISP: A Multi-lingual Multi-accent Dataset for Speaker Profiling

ICASSP 2021accepted

Many commercial and forensic applications of speech demand the extraction of information about the speaker characteristics, which falls into the broad category of speaker profiling. The speaker characteristics needed for profiling include physical traits of the speaker like height, age, and gender o…

Cited by 0SourceScholar
2021

Representation Learning for Speech Recognition Using Feedback Based Relevance Weighting

ICASSP 2021accepted

In this work, we propose an acoustic embedding based approach for representation learning in speech recognition. The proposed approach involves two stages comprising of acoustic filterbank learning from raw waveform, followed by modulation filterbank learning. In each stage, a relevance weighting op…

Cited by 0SourceScholar
2020

3-D Acoustic Modeling for Far-Field Multi-Channel Speech Recognition

ICASSP 2020accepted

The conventional approach to automatic speech recognition in multichannel reverberant conditions involves a beamforming based enhancement of the multi-channel speech signal followed by a single channel neural acoustic model. In this paper, we propose to model the multi-channel signal directly using…

Cited by 0SourceScholar
2020

Improving Voice Separation by Incorporating End-To-End Speech Recognition

ICASSP 2020accepted

Despite recent advances in voice separation methods, many challenges remain in realistic scenarios such as noisy recording and the limits of available data. In this work, we propose to explicitly incorporate the phonetic and linguistic nature of speech by taking a transfer learning approach using an…

Cited by 0SourceScholar
2020

On The Impact of Language Familiarity in Talker Change Detection

ICASSP 2020accepted

The ability to detect talker changes when listening to conversational speech is fundamental to perception and understanding of multi-talker speech. In this paper, we propose an experimental paradigm to provide insights on the impact of language familiarity on talker change detection. Two multi-talke…

Cited by 0SourceScholar
2020

Unsupervised Neural Mask Estimator for Generalized Eigen-Value Beamforming Based Asr

ICASSP 2020accepted

The state-of-art methods for acoustic beamforming in multi-channel ASR are based on a neural mask estimator that predicts the presence of speech and noise. These models are trained using a paired corpus of clean and noisy recordings (teacher model). In this paper, we attempt to move away from the re…

Cited by 0SourceScholar
2019

A Deep Neural Network Based End to End Model for Joint Height and Age Estimation from Short Duration Speech

ICASSP 2019accepted

Automatic height and age prediction of a speaker has a wide variety of applications in speaker profiling, forensics etc. Often in such applications only a few seconds of speech data is available to reliably estimate the speaker parameters. Traditionally, age and height were predicted separately usin…

Cited by 0SourceScholar
2019

Analyzing Human Reaction Time for Talker Change Detection

ICASSP 2019accepted

The ability to detect a change in the input is an essential aspect of perception. In speech communication, we use this ability to identify "talker changes" when listening to conversational speech (such as, audio podcasts). In this paper, we propose to improve our understanding about how fast listene…

Cited by 0SourceScholar
2019

End-to-end Language Recognition Using Attention Based Hierarchical Gated Recurrent Unit Models

ICASSP 2019accepted

The task of automatic language identification (LID) involving multiple dialects of the same language family on short speech recordings is a challenging problem. This can be further complicated for short-duration audio snippets in the presence of noise sources. In these scenarios, the identity of the…

Cited by 0SourceScholar
2019

The Leap Speaker Recognition System for NIST SRE 2018 Challenge

ICASSP 2019accepted

The NIST Speaker Recognition Evaluation (SRE) 2018 challenge comprises an open evaluation of the text independent speaker verification task. This paper summarizes the LEAP speaker verification systems submitted to the NIST SRE 2018. For all the speaker verification approaches, the front-end feature…

Cited by 0SourceScholar
2018

Enhancement and Analysis of Conversational Speech: JSALT 2017

ICASSP 2018accepted

Automatic speech recognition is more and more widely and effectively used. Nevertheless, in some automatic speech analysis tasks the state of the art is surprisingly poor. One of these is “diarization”, the task of determining who spoke when. Diarization is key to processing meeting audio and clinic…

Cited by 0SourceScholar
2018

Leveraging LSTM Models for Overlap Detection in Multi-Party Meetings

ICASSP 2018accepted

The detection of overlapping speech segments is of key importance in speech applications involving analysis of multi-party conversations. The detection problem is challenging because overlapping speech segments are typically captured as short speech utterances far-field microphone recordings. In thi…

Cited by 0SourceScholar
2017

Factor analysis methods for joint speaker verification and spoof detection

ICASSP 2017accepted

The performance of a speaker verification system is severely degraded by spoofing attacks generated from artificial speech synthesizers. Recently, several approaches have been proposed for classifying natural and synthetic speech (spoof detection) which can be used in conjunction with a speaker veri…

Cited by 0SourceScholar
2016

Speaker age estimation on conversational telephone speech using senone posterior based i-vectors

ICASSP 2016accepted

Automatic age estimation from speech has a variety of applications including natural human-computer interaction, targeted advertising, customer-agent pairing in call centers, and forensics, to mention a few. Recently, the use of i-vectors has shown promise for automatic age estimation. In this paper…

Cited by 0SourceScholar
2015

Nearest neighbor discriminant analysis for language recognition

ICASSP 2015accepted

Many state-of-the-art i-vector based voice biometric systems use linear discriminant analysis (LDA) as a post-processing stage to increase the computational efficiency in the back-end via dimensionality reduction, as well as annihilate the undesired (noisy) directions in the total variability subspa…

Cited by 0SourceScholar