← Search

Beena Ahmed

9 accepted papers

2025

Evidential Neural GPLDA: A Novel Approach to Quantify Prediction Uncertainty in Speaker Verification Systems

ICASSP 2025accepted

The uncertainty of an automatic speaker verification (ASV) system is typically estimated using its overall accuracy. However it fails to express "when" the system is uncertain in a predictive and case-by-case manner. Also, prior to interpreting each prediction made by ASV systems, there is a need to…

Cited by 0SourceScholar
2025

Improved Out-of-domain Detection in VAE Latent Spaces with Boundary-driven Regularisation

ICASSP 2025accepted

In out-of-domain (OOD) detection tasks, encoding the actual data into a suitable latent space could be beneficial since it may facilitate measurement of the spatial relationship between in-domain (IND) and OOD data. However, any such mapping of data to a latent space carries the risk that some OOD p…

Cited by 0SourceScholar
2025

Multi-Class Dementia Detection Using Acoustic Features - ICASSP-2025 PROCESS Challenge

ICASSP 2025accepted

This paper describes our best-performing submission for the ICASSP-2025 Signal Processing Grand Challenge PROCESS, focused on the classification of speech into 3 groups - Healthy, Mild Cognitive Impairment (MCI), and Dementia - using three speech tasks in English. Our approach was aligned with the a…

Cited by 0SourceScholar
2025

Rethinking Mamba in Speech Processing by Self-Supervised Models

ICASSP 2025accepted

The Mamba-based model has demonstrated outstanding performance across tasks in computer vision, natural language processing, and speech processing. However, in the realm of speech processing, the Mamba-based model’s performance varies across different tasks. For instance, in tasks such as speech enh…

Cited by 0SourceScholar
2025

SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information

ACL 2025finding

Large Language Models (LLMs) have been increasingly adopted for health-related tasks, yet their performance in depression detection remains limited when relying solely on text input. While Retrieval-Augmented Generation (RAG) typically enhances LLM capabilities, our experiments indicate that traditi…

Cited by 0SourcePDFScholar
2024

A Probability Gradient Based Approach for Sampling Boundaries of In-Domain Data

ICASSP 2024accepted

In machine learning applications, it is desirable to distinguish between in-domain and out-of-domain data. However, in most cases, only in-domain data is available and consequently identifying the ‘boundary’ between in-domain and out-of-domain is a significant challenge. In this paper we present a n…

Cited by 0SourceScholar
2024

Variational Connectionist Temporal Classification for Order-Preserving Sequence Modeling

ICASSP 2024accepted

Connectionist temporal classification (CTC) is commonly adopted for sequence modeling tasks like speech recognition, where it is necessary to preserve order between the input and target sequences. However, CTC is only applied to deterministic sequence models, where the latent space is discontinuous…

Cited by 0SourceScholar
2024

When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection

EMNLP 2024main

Depression is a critical concern in global mental health, prompting extensive research into AI-based detection methods. Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in healthcare applications. However, the application of LLMs in the identification and a…

Cited by 11SourcePDFScholar
2016

Classification of bisyllabic lexical stress patterns in disordered speech using deep learning

ICASSP 2016accepted

Technology-based therapy tools can be of great benefit to children with developmental speech disabilities as they typically require sustained practice with a speech therapist for several years. Towards this aim, over the past 4 years we have developed speech processing tools to automatically detect…

Cited by 0SourceScholar