← Search

Eliathamby Ambikairajah

16 accepted papers

2025

Blind Estimation of Sub-band Acoustic Parameters from Ambisonics Recordings using Spectro-Spatial Covariance Features

ICASSP 2025accepted

Estimating frequency-varying acoustic parameters is essential for enhancing immersive perception in realistic spatial audio creation. In this paper, we propose a unified framework that blindly estimates reverberation time (T60), direct-to-reverberant ratio (DRR), and clarity (C50) across 10 frequenc…

Cited by 0SourceScholar
2024

An Empirical Study on the Impact of Positional Encoding in Transformer-Based Monaural Speech Enhancement

ICASSP 2024accepted

Transformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to enable Transformers to distinguish the order of elements in a sequence. However, it remains unclear how positional encod…

Cited by 0SourceScholar
2023

Constrained Dynamical Neural ODE for Time Series Modelling: A Case Study on Continuous Emotion Prediction

ICASSP 2023accepted

weA number of machine learning applications involve time series prediction, and in some cases additional information about dynamical constraints on the target time series may be available. For instance, it might be known that the desired quantity cannot change faster than some rate or that the rate…

Cited by 0SourceScholar
2022

A Novel Sequential Monte Carlo Framework for Predicting Ambiguous Emotion States

ICASSP 2022accepted

When continuous emotion labelling of natural (non-acted) data is desired, it is typically collected from multiple annotators. However, most automatic emotion recognition systems trained on such data ignore disagreement between annotators and only models the average rating, despite the observation th…

Cited by 0SourceScholar
2020

Adversarial Multi-Task Learning for Speaker Normalization in Replay Detection

ICASSP 2020accepted

Spoofing detection algorithms in voice biometrics are adversely affected by differences in the speech characteristics of the various target users. In this paper, we propose a novel speaker normalisation technique that employs adversarial multi-task learning to compensate for this speaker variability…

Cited by 0SourceScholar
2020

Cochlear Signal Processing: A Platform for Learning the Fundamentals of Digital Signal Processing

ICASSP 2020accepted

The first digital signal processing course in most electrical engineering programmes around the world tends to be a significant jump in abstraction for most students. This is a consequence of them being introduced to a large number of mathematical concepts with insufficient time to consolidate the i…

Cited by 0SourceScholar
2019

Auditory Inspired Spatial Differentiation for Replay Spoofing Attack Detection

ICASSP 2019accepted

The security of Automatic Speaker Verification systems is greatly threatened by spoofing attacks of various kinds. Among them, replay attacks are noteworthy due to the ease with which they can be employed. Most countermeasures for replay attacks use subband features based on parallel filter banks. T…

Cited by 0SourceScholar
2019

Evaluation Measures for Depression Prediction and Affective Computing

ICASSP 2019accepted

A variety of evaluation measures are being used to validate systems in depression prediction and affective computing. Among them, the most common measures focus on the error between the ground truth and predictions. However, when the ground truth is ordinal such as in psychiatric scores, ranking inf…

Cited by 0SourceScholar
2019

Phoneme Specific Modelling and Scoring Techniques for Anti Spoofing System

ICASSP 2019accepted

Replay attack refers to the use of recorded speech in an attempt to spoof an automatic speaker verification system and the development of countermeasures that can detect these attacks is an active area of research. This paper investigates the effect of phoneme specific information on replay attack d…

Cited by 0SourceScholar
2019

Transmission Line Cochlear Model Based AM-FM Features for Replay Attack Detection

ICASSP 2019accepted

This paper focuses on providing a countermeasure to replay attack which is the simplest and more accessible form of attack used to spoof automatic speaker verification systems. Specifically, it proposes the use of the transmission line cochlear model, which resembles the human cochlea more accuratel…

Cited by 13SourceScholar
2018

Dynamic Multi-Rater Gaussian Mixture Regression Incorporating Temporal Dependencies of Emotion Uncertainty Using Kalman Filters

ICASSP 2018accepted

Predicting continuous emotion in terms of affective attributes has mainly been focused on hard labels, which ignored the ambiguity of recognizing certain emotions. This ambiguity may result in high inter-rater variability and in turn causes varying prediction uncertainty with time. Based on the assu…

Cited by 0SourceScholar
2018

End-to-End Hierarchical Language Identification System

ICASSP 2018accepted

Recently, hierarchical language identification systems have shown significant improvement over single level systems in both closed and open set language identification tasks. However, developing such a system requires the features and classifier selection at each node in the hierarchical structure t…

Cited by 0SourceScholar
2018

Factorized Hidden Variability Learning for Adaptation of Short Duration Language Identification Models

ICASSP 2018accepted

Bidirectional long short term memory (BLSTM) recurrent neural networks (RNNs) have recently outperformed other state-of-the-art approaches, such as i-vector and deep neural networks (DNNs) in automatic language identification (LID), particularly when testing with very short utterances (`3s). Mismatc…

Cited by 6SourceScholar
2018

Speaker-Phonetic Vector Estimation for Short Duration Speaker Verification

ICASSP 2018accepted

Phonetic variability is one of the primary challenges in short duration speaker verification. This paper proposes a novel method that modifies the standard normal distribution prior in the total variability model to use a mixture of Gaussians as the prior distribution. The proposed speaker-phonetic…

Cited by 0SourceScholar
2017

Salience based lexical features for emotion recognition

ICASSP 2017accepted

In this paper we focus on the usefulness of verbal events for speech based emotion recognition. In particular, the use of phoneme sequences to encode verbal cues related to the expression of emotions is proposed and lexical features based on these phoneme sequences are introduced for use in automati…

Cited by 0SourceScholar
2016

A hierarchical framework for language identification

ICASSP 2016accepted

Most current language recognition systems model different levels of information such as acoustic, prosodic, phonotactic, etc. independently and combine the model likelihoods in order to make a decision. However, these are single level systems that treat all languages identically and hence incapable…

Cited by 0SourceScholar