← Search

Vidhyasaharan Sethu

19 accepted papers

2025

AER-LLM: Ambiguity-aware Emotion Recognition Leveraging Large Language Models

ICASSP 2025accepted

Recent advancements in Large Language Models (LLMs) have demonstrated great success in many Natural Language Processing (NLP) tasks. In addition to their cognitive intelligence, exploring their capabilities in emotional intelligence is also crucial, as it enables more natural and empathetic conversa…

Cited by 0SourceScholar
2025

Blind Estimation of Sub-band Acoustic Parameters from Ambisonics Recordings using Spectro-Spatial Covariance Features

ICASSP 2025accepted

Estimating frequency-varying acoustic parameters is essential for enhancing immersive perception in realistic spatial audio creation. In this paper, we propose a unified framework that blindly estimates reverberation time (T60), direct-to-reverberant ratio (DRR), and clarity (C50) across 10 frequenc…

Cited by 0SourceScholar
2025

Evidential Neural GPLDA: A Novel Approach to Quantify Prediction Uncertainty in Speaker Verification Systems

ICASSP 2025accepted

The uncertainty of an automatic speaker verification (ASV) system is typically estimated using its overall accuracy. However it fails to express "when" the system is uncertain in a predictive and case-by-case manner. Also, prior to interpreting each prediction made by ASV systems, there is a need to…

Cited by 0SourceScholar
2025

Improved Out-of-domain Detection in VAE Latent Spaces with Boundary-driven Regularisation

ICASSP 2025accepted

In out-of-domain (OOD) detection tasks, encoding the actual data into a suitable latent space could be beneficial since it may facilitate measurement of the spatial relationship between in-domain (IND) and OOD data. However, any such mapping of data to a latent space carries the risk that some OOD p…

Cited by 0SourceScholar
2024

A Probability Gradient Based Approach for Sampling Boundaries of In-Domain Data

ICASSP 2024accepted

In machine learning applications, it is desirable to distinguish between in-domain and out-of-domain data. However, in most cases, only in-domain data is available and consequently identifying the ‘boundary’ between in-domain and out-of-domain is a significant challenge. In this paper we present a n…

Cited by 0SourceScholar
2024

Variational Connectionist Temporal Classification for Order-Preserving Sequence Modeling

ICASSP 2024accepted

Connectionist temporal classification (CTC) is commonly adopted for sequence modeling tasks like speech recognition, where it is necessary to preserve order between the input and target sequences. However, CTC is only applied to deterministic sequence models, where the latent space is discontinuous…

Cited by 0SourceScholar
2023

Constrained Dynamical Neural ODE for Time Series Modelling: A Case Study on Continuous Emotion Prediction

ICASSP 2023accepted

weA number of machine learning applications involve time series prediction, and in some cases additional information about dynamical constraints on the target time series may be available. For instance, it might be known that the desired quantity cannot change faster than some rate or that the rate…

Cited by 0SourceScholar
2022

A Novel Sequential Monte Carlo Framework for Predicting Ambiguous Emotion States

ICASSP 2022accepted

When continuous emotion labelling of natural (non-acted) data is desired, it is typically collected from multiple annotators. However, most automatic emotion recognition systems trained on such data ignore disagreement between annotators and only models the average rating, despite the observation th…

Cited by 0SourceScholar
2020

Adversarial Multi-Task Learning for Speaker Normalization in Replay Detection

ICASSP 2020accepted

Spoofing detection algorithms in voice biometrics are adversely affected by differences in the speech characteristics of the various target users. In this paper, we propose a novel speaker normalisation technique that employs adversarial multi-task learning to compensate for this speaker variability…

Cited by 0SourceScholar
2020

Cochlear Signal Processing: A Platform for Learning the Fundamentals of Digital Signal Processing

ICASSP 2020accepted

The first digital signal processing course in most electrical engineering programmes around the world tends to be a significant jump in abstraction for most students. This is a consequence of them being introduced to a large number of mathematical concepts with insufficient time to consolidate the i…

Cited by 0SourceScholar
2019

Auditory Inspired Spatial Differentiation for Replay Spoofing Attack Detection

ICASSP 2019accepted

The security of Automatic Speaker Verification systems is greatly threatened by spoofing attacks of various kinds. Among them, replay attacks are noteworthy due to the ease with which they can be employed. Most countermeasures for replay attacks use subband features based on parallel filter banks. T…

Cited by 0SourceScholar
2019

Phoneme Specific Modelling and Scoring Techniques for Anti Spoofing System

ICASSP 2019accepted

Replay attack refers to the use of recorded speech in an attempt to spoof an automatic speaker verification system and the development of countermeasures that can detect these attacks is an active area of research. This paper investigates the effect of phoneme specific information on replay attack d…

Cited by 0SourceScholar
2018

Dynamic Multi-Rater Gaussian Mixture Regression Incorporating Temporal Dependencies of Emotion Uncertainty Using Kalman Filters

ICASSP 2018accepted

Predicting continuous emotion in terms of affective attributes has mainly been focused on hard labels, which ignored the ambiguity of recognizing certain emotions. This ambiguity may result in high inter-rater variability and in turn causes varying prediction uncertainty with time. Based on the assu…

Cited by 0SourceScholar
2018

End-to-End Hierarchical Language Identification System

ICASSP 2018accepted

Recently, hierarchical language identification systems have shown significant improvement over single level systems in both closed and open set language identification tasks. However, developing such a system requires the features and classifier selection at each node in the hierarchical structure t…

Cited by 0SourceScholar
2018

Factorized Hidden Variability Learning for Adaptation of Short Duration Language Identification Models

ICASSP 2018accepted

Bidirectional long short term memory (BLSTM) recurrent neural networks (RNNs) have recently outperformed other state-of-the-art approaches, such as i-vector and deep neural networks (DNNs) in automatic language identification (LID), particularly when testing with very short utterances (`3s). Mismatc…

Cited by 6SourceScholar
2018

Speaker-Phonetic Vector Estimation for Short Duration Speaker Verification

ICASSP 2018accepted

Phonetic variability is one of the primary challenges in short duration speaker verification. This paper proposes a novel method that modifies the standard normal distribution prior in the total variability model to use a mixture of Gaussians as the prior distribution. The proposed speaker-phonetic…

Cited by 0SourceScholar
2017

Salience based lexical features for emotion recognition

ICASSP 2017accepted

In this paper we focus on the usefulness of verbal events for speech based emotion recognition. In particular, the use of phoneme sequences to encode verbal cues related to the expression of emotions is proposed and lexical features based on these phoneme sequences are introduced for use in automati…

Cited by 0SourceScholar
2016

A hierarchical framework for language identification

ICASSP 2016accepted

Most current language recognition systems model different levels of information such as acoustic, prosodic, phonotactic, etc. independently and combine the model likelihoods in order to make a decision. However, these are single level systems that treat all languages identically and hence incapable…

Cited by 0SourceScholar
2015

Weighted pairwise Gaussian likelihood regression for depression score prediction

ICASSP 2015accepted

This paper presents a technique in which feature vectors are mapped onto ordinal ranges of clinical depression scores using weighted pairwise Gaussians. The position of a test vector with respect to these partitions is used to perform depression score prediction. Results found on a set of spectral a…

Cited by 0SourceScholar